keyword filters miss the pattern. a score is not a reason. and someone else's taxonomy is not your policy.
| 1 | flash | rules & cache | checks explicit policy rules and matching previous decisions. unresolved messages move to echo. |
| 2 | echo | content classification | evaluates supported moderation categories. results that need further evaluation move to deep. |
| 3 | deep | policy reasoning | examines ambiguous content against your policy and the context you provide. returns a decision with a policy reference, or flags the message for review. |
result: allow, block, or review, with the stage used and the policy version. today flash runs rules, echo is model classification, deep reads a demo policy. planned: a cache in flash, your own policy in deep.
planning assumptions, stated plainly, and the network that carries them. measured results replace the targets as reports are published below.
stage execution, not complete request time. targets, not measurements.
recall target > 90% across categories. targets, not measurements.
nothing is published until it comes from a reproducible run. the first report ships with its method attached.
measured on a fixed reviewed set and a fresh challenge set, cold and warm, with confidence intervals. today: flash runs rules, echo runs gpt-oss-safeguard-20b, deep runs gpt-oss-safeguard-120b. a stage is final at allow ≥ 0.85 or block ≥ 0.90.
hard cases become new rules and sharper rubric wording. every change passes the evals before it ships.
hard cases become new flash rules and sharper rubric wording.
a fixed set and a fresh challenge set. no category may lose recall.
every promotion is a version bump. rollback is one call.
one email when keys are ready. nothing else.