Skip to main content
ScienceVerse

← All experiments

No supervision neededSafety review pendingAges 8–16

Biased Training Cards

Sign in to save this experiment.

Safety first

No supervision needed

No adult supervision needed: paper cards only.

Protective equipment and “do not substitute” warnings have not been recorded for this experiment yet. They are added when its safety review is completed; until then, follow the supervision notes above.

Disposal: Recycle the cards, or keep them to repeat the activity.

Materials

  • About 20 small pieces of card
  • Coloured pens or pencils
  • A partner

Steps

  1. Make 10 training cards. Draw 5 red apples on 5 cards, and a banana, an orange, a lemon, a pear and a bunch of grapes on the other 5.
  2. Make 6 test cards: a green apple, a yellow apple, a red tomato, a red ball, a cherry and another red apple.
  3. Your partner is the 'learner'. Show them only the training cards, telling them which are apples, and ask them to write a rule for spotting an apple.
  4. Now show the test cards one at a time. The learner must use only their written rule to say 'apple' or 'not apple'.
  5. Count the mistakes. Then add a green apple and a yellow apple to the training cards, let the learner rewrite the rule, and test again.

What you should see

  • The first rule is usually about colour, such as 'red and round'.
  • That rule calls the green and yellow apples 'not apple' and calls the red tomato and red ball 'apple'.
  • After the training cards include apples of other colours, the rule changes and there are fewer mistakes.

Why it works

A machine-learning model learns patterns from the examples it is trained on. If those examples do not reflect the variety of the real world, the patterns it learns are wrong for the cases it never saw. This is called selection bias; when part of the real world is missing from the data, it is called coverage bias.

Your learner did exactly what a model does: every apple they saw was red, so 'red' looked like the most useful clue. The fix was not a cleverer learner but better, more representative training data.

Learn the ideas behind it

Sources