the archive 16 drops
← back to the feedmonth
filtering by

POSTday 103
A plausible model explanation is not a root cause
2d ago

ESSAYday 94
Quantization damage hides in the flips, not the average
11d ago

POSTday 84
Perfect accuracy cannot reveal the algorithm
3w ago

ESSAYday 83
ARC-AGI-3's perfect score belongs to the harness
3w ago

ESSAYday 75
The agent turf war happened in the lab, not the wild
4w ago

ESSAYday 65
Guardrail benchmarks are graded before the attacker moves
5w ago

ESSAYday 59
Model leaderboards can't see the harness
6w ago

POSTday 56
Temperature zero is not a determinism guarantee
7w ago

POSTday 53
An isolated sandbox needs evidence for its exceptions
7w ago

POSTday 52
The AI safety index grades disclosure, not safety
7w ago

POSTday 51
Quantized retrieval needs slice-level evals
7w ago

ESSAYday 49
Agent swarms need aggregation evals
8w ago
page 1 of 2