Back to feed
New livenerf tool tests AI quality post-release

Open benchmark livenerf checks AI model quality after release

Developer ninjahawk launched the livenerf project — an open benchmark to test AI model quality degradation after release. The tool is designed for objective assessment of user complaints regarding reduced effectiveness of systems such as Claude Opus 5.5.

Testing methodology is based on strict task selection and control of environment parameters.

From an initial set of 2,336 questions (GPQA Diamond, MMLU-Pro, olympiad mathematics), 78 tasks were selected where the model answers correctly only occasionally. The test panel runs daily for 30 days, with the first 10 serving as a baseline. Prompts are frozen, the Claude Code version is fixed, and responses are verified by exact match without an AI judge.

To exclude platform changes' influence, parallel control of the Opus 5 model is conducted. If both models show a drop, it indicates an infrastructure issue rather than a problem with the Opus 5.5 version itself.

The key indicator of potential deterioration is the number of output tokens. A decrease in this metric may signal that the model is "thinking less," which often precedes a drop in overall answer accuracy.

The technical implementation uses Inspect, an open framework from the UK's AI safety institute. Testing is available via a standard Max subscription without needing an API key. Initial results are expected around October 24, following six days of data collection.

4.6K views

More from this channel Black Triangle Channel Telegram

Similar in this category Technology