Ali Toygar Abak PRO
phionyx
AI & ML interests
AI governance, AI safety, multi-agent systems, reproducible LLM evaluation, agent protocols, open-source AI, local inference, and trustworthy AI
Recent Activity
liked a dataset about 7 hours ago
evaleval/EEE_datastore repliedto their post about 10 hours ago
AI evaluation results can look more certain than they actually are.
A run finishes. A dashboard shows PASS. A score gets copied into another system. A few steps later, it may no longer be clear which metric produced it, what was skipped, what population was actually tested, or whether that PASS came from the evaluator at all.
That is the problem behind a new individual IETF Internet-Draft I published today: Claim-Preserving Exchange of AI Evaluation Evidence.
The basic idea is simple:
moving evidence should not make the claim stronger than the evidence itself.
Technical review and counterexamples are very welcome:
https://datatracker.ietf.org/doc/html/draft-abak-ai-evaluation-claim-preservation published an article about 10 hours ago
The Dashboard Says PASS. What Exactly Passed?Organizations
None yet