02 / AI-SAFETY EVALS
Evaluation harnesses for AI systems
We build evaluation harnesses for AI systems using the same planted-twin discipline as the smart-contract work. A language-conditioned detection-rate eval, a regression harness, or a ground-truth benchmark built from planted-bug twins. The eval harness measures where an AI system's accuracy degrades and surfaces the failure mode in a reproducible, CI-runnable form.
Scope: eval harness design and build; planted-bug ground truth generation; detection-rate measurement across language, model, or domain conditions. We deliver a harness and a proof register, same format as the contract work. Not a standalone AI audit service.
Proof: apart-global-south-lost-in-translation: language-conditioned detection-rate eval, EN / ES / PT / CS, Atlas planted-bug twins as ground truth. Apart Global South track submission. Research artifact and eval harness.