← Writing

Collection

Published cyber benchmark results

I collected published results from 19 cyber benchmarks for 2026 models with an Artificial Analysis intelligence index above 50.

Open the full benchmark table 19 benchmarks · Fileverse

Turns out it is surprisingly difficult to compare these models despite the number of available evals.

Most lack published results on the same benchmark. Even the most popular like DeepsecBench have gaps.

5 of 19 benchmarks had published results for at least 10 models in the collection.

The collection was made by whipping ChatGPT Pro many times. Enjoy and comment with your experience (or my misses).