The question
A fluent answer can hide missing evidence. I wanted a small, inspectable way to measure when a system answers correctly, invents an answer, or declines a question it could have answered.
AbstainBench contains 30 questions across science, math, geography, history, and fictional missing-evidence prompts.
What I built
I built the browser interface, fixed question set, scoring, and JSON and Markdown exports. The evaluation records answer accuracy, abstention accuracy, hallucinations, false abstentions, and results by category.
Two modes serve different purposes. A transparent offline baseline checks the labels, scoring, and export pipeline. An optional local model uses WebLLM and WebGPU to run quantized Llama 3.2 1B in the browser.
Try it
Start with the offline baseline to inspect how the evaluation works. It does not need an API key or model download.
The optional local-model mode downloads model weights on first use. It requires a compatible WebGPU browser and suitable hardware. Treat the resulting scores as a small experiment.
The important limitation
The offline baseline uses the known labels. Its score is not model performance. Thirty questions are also too few to support a general claim about a model's reliability.
A stronger evaluation would add independently reviewed labels, more models, and confidence intervals across repeated runs. The value of this version is that its assumptions and outputs are easy to inspect.
