Risk and safety

Sycophancy

A model's tendency to agree with the user, praise their ideas or change a correct answer when pushed back on. It is partly a side effect of training on human approval. Ask for criticism explicitly, and be wary when a model instantly agrees it was wrong.


Related terms

AI incident

An event where an AI system causes or nearly causes harm, such as a chatbot giving customers wrong refund terms or an …

AI safety

The field concerned with preventing harm from AI systems, from everyday failures such as biased or false outputs to se…

AI text detector

A tool that claims to tell whether text was written by AI. They are unreliable: they produce false positives, particul…

Adversarial example

An input subtly altered to fool a model, such as an image with changes invisible to people that makes a classifier see…

Alignment

Making an AI system's behaviour match the intentions and values of the people it serves, including in situations its d…

Automation bias

The human tendency to over-trust an automated system's output and stop checking it, especially when it is usually righ…

Beyond the definition

Knowing the word is the easy part

A short assessment scores you across five skill areas and builds a path through 18 modules, skipping whatever you already know. Free, no card.

Find your level