Skip to content
Safety

Constitutional AI

A training approach where AI models are given a set of principles (a 'constitution') and learn to self-critique and revise their outputs to comply with those principles. Reduces reliance on human feedback for safety alignment.

Why it matters

Teaching AI to behave well usually means paying people to rate thousands of responses, which is slow and hard to scale. Constitutional AI gives the model a written set of principles and lets it critique and fix its own answers against them. This makes safety training more consistent and transparent, since the guiding rules are actually written down rather than buried in scattered human judgments.

In practice

A model drafts a reply that's technically correct but rude. Under Constitutional AI, it checks its draft against a principle like "be helpful and respectful," notices the tone problem, and rewrites the answer to be polite while keeping the useful content. This self-review happens during training, so by the time you use the model, it has already learned to lean toward responses that follow those principles.

Related terms

Put Constitutional AI into practice

Access 800+ AI models and 70+ tools through Vincony — start free with 100 credits.