LLMs want to guess.
Production systems refuse.

Automated research that traces every claim to its source, catches contradictions in code rather than in a prompt, and stops before writing a conclusion it cannot support.

Watch it refuse

Raw evidence

Execution

Nothing yet.
Six sources go in. Watch what comes out.

Edit the evidence and it will conclude. The threshold is arithmetic, not a prompt.