evaluate
To evaluate an AI system means to systematically test how well it performs, using set questions, tasks, or criteria, rather than just trying it out informally. This can cover accuracy, safety, tone, or how well it handles specific tasks, and is used both by developers building AI tools and by people choosing or configuring an AI agent. Evaluation results help decide whether a model or agent is good enough to rely on for a given purpose.
Related terms
A changelog page is a running list, usually in a product's documentation site, of updates, new features, fixes and oth…
connectorsConnectors are pre-built links that let an AI tool or assistant read data from, or take actions in, other apps and ser…
documentation indexA single file, often called llms.txt, that lists all the pages in a product's documentation so that an AI tool or chat…
organizationAn organization (or workspace) is the shared account that groups together the people, projects, billing, and security …
release notesRelease notes are a running log that an AI provider publishes each time it updates a tool or model, listing what's new…
responsesThe Responses API is OpenAI's interface for sending input to a model and getting back its output, including text, tool…
Knowing the word is the easy part
A short assessment scores you across five skill areas and builds a path through 18 modules, skipping whatever you already know. Free, no card.
Find your level