What changed
MLCommons tagged MLPerf Client v2.0 on August 18, 2026, adding an Agentic AI Benchmarking category whose Software Engineering (SWE) and Data Analyst tasks track end-to-end performance including tool execution time 1. The LLM lineup moved with it, and the table below carries the detail 1. A new image generation category features Flux 2 Klein 4B, experimental 1.
What the score now contains
The measurement boundary moved off the model and onto the whole local stack. The agentic benchmarks time tool execution alongside inference, so the score belongs to the device plus whatever runs the SWE and Data Analyst tasks 1. The notes name both task families and name neither the agent harness nor the tools behind them 1. A year earlier the 1.0 writeup named Llama 2 7B Chat and Phi 3.5 Mini Instruct among the models it tested 2; v2.0 removes Phi 3.5 and moves Phi 4 Reasoning 14B into the extended category 1.
| Model | Named in the 1.0 writeup, a year earlier 2 | Status in v2.0 1 |
|---|---|---|
| Llama 3.1 8B Instruct | tested | mandatory base benchmark |
| Phi 3.5 Mini Instruct | tested | removed |
| Phi 4 Reasoning 14B | experimental | moved to the extended category |
| Llama 2 7B Chat | tested | not named in the v2.0 notes |
| Phi 4 Mini Instruct | not among the models that writeup names | mandatory base benchmark |
| Qwen 3 8B | not among the models that writeup names | experimental test |
Impact on your team
This lands on anyone buying or publishing AI PC numbers, or timing an agent workload. Re-baseline on v2.0 rather than diffing a v1.x archive: Phi 3.5 is gone from the base set 1, and the lineup that writeup documented 2 is not v2.0’s 1. When a vendor quotes an agentic score, demand the task family and the tool layer behind it before you compare two devices 1. Re-run it yourself or it is not your number.