DEV Community

#benchmark

Posts

👋 Sign in for the ability to sort posts by relevant, latest, or top.
I Tested Q4_K_M vs MXFP4 on the Same Laptop — The Supposedly-Faster New Format Lost

I Tested Q4_K_M vs MXFP4 on the Same Laptop — The Supposedly-Faster New Format Lost

Comments
6 min read
āđ€āļĄāļ·āđˆāļ­ Benchmark āđ‚āļāļŦāļāļ„āļļāļ“, SWE-Bench ProMax āļāļąāļšāļ„āļ°āđāļ™āļ™āļˆāļĢāļīāļ‡āļ—āļĩāđˆāđ‚āļĄāđ€āļ”āļĨāđ€āļāđˆāļ‡āļŠāļļāļ”āļ—āļģāđ„āļ”āđ‰āđāļ„āđˆ 41.2%

āđ€āļĄāļ·āđˆāļ­ Benchmark āđ‚āļāļŦāļāļ„āļļāļ“, SWE-Bench ProMax āļāļąāļšāļ„āļ°āđāļ™āļ™āļˆāļĢāļīāļ‡āļ—āļĩāđˆāđ‚āļĄāđ€āļ”āļĨāđ€āļāđˆāļ‡āļŠāļļāļ”āļ—āļģāđ„āļ”āđ‰āđāļ„āđˆ 41.2%

Comments
2 min read
āļŠāđˆāļ­āļ‡āļ§āđˆāļēāļ‡ 0.3% āđāļ•āđˆāļĢāļēāļ„āļēāļ•āđˆāļēāļ‡ 2 āđ€āļ—āđˆāļē, āļ­āđˆāļēāļ™āļ•āļēāļĢāļēāļ‡ Terminal-Bench 4.0 āđƒāļŦāđ‰āđ€āļ›āđ‡āļ™

āļŠāđˆāļ­āļ‡āļ§āđˆāļēāļ‡ 0.3% āđāļ•āđˆāļĢāļēāļ„āļēāļ•āđˆāļēāļ‡ 2 āđ€āļ—āđˆāļē, āļ­āđˆāļēāļ™āļ•āļēāļĢāļēāļ‡ Terminal-Bench 4.0 āđƒāļŦāđ‰āđ€āļ›āđ‡āļ™

Comments
2 min read
Benchmarking Real-Time Voice AI APIs: Cartesia vs Deepgram vs ElevenLabs (2026)

Benchmarking Real-Time Voice AI APIs: Cartesia vs Deepgram vs ElevenLabs (2026)

Comments
1 min read
I built a benchmark for AI companion apps because every "best AI girlfriend" list is affiliate spam

I built a benchmark for AI companion apps because every "best AI girlfriend" list is affiliate spam

Comments 1
5 min read
A Benchmark Is Only as Honest as Its Harness

A Benchmark Is Only as Honest as Its Harness

Comments
4 min read
Extraction APIs invented values for 17% of the fields that aren't in the document. Ours included.

Extraction APIs invented values for 17% of the fields that aren't in the document. Ours included.

Comments
12 min read
A Small Transformer Trained in 1.5 Hours Beat Many LLMs on ARC

A Small Transformer Trained in 1.5 Hours Beat Many LLMs on ARC

Comments
5 min read
I Ran 3 Open-Weight LLMs Head-to-Head on a 24GB Mac — One Was 3x Faster

I Ran 3 Open-Weight LLMs Head-to-Head on a 24GB Mac — One Was 3x Faster

2
Comments 4
5 min read
Benchmark a Free AI Coding Tier on a Cold Server

Benchmark a Free AI Coding Tier on a Cold Server

Comments
4 min read
AxonASP vs. Native IIS ASP: Performance Benchmarks and Engine Architecture

AxonASP vs. Native IIS ASP: Performance Benchmarks and Engine Architecture

Comments
2 min read
We ran 160 agent tasks across two frameworks. The frameworks tied. Then we changed the model.

We ran 160 agent tasks across two frameworks. The frameworks tied. Then we changed the model.

Comments
4 min read
A LongMemEval-S number you can reproduce

A LongMemEval-S number you can reproduce

Comments 4
6 min read
I Tested GLM-5.3-Flash and Qwen3.8-Flash on 24 Real Tasks

I Tested GLM-5.3-Flash and Qwen3.8-Flash on 24 Real Tasks

Comments
6 min read
Twenty Prompts, One Compiler: A Free AI Server's Failure Matrix

Twenty Prompts, One Compiler: A Free AI Server's Failure Matrix

Comments
5 min read
👋 Sign in for the ability to sort posts by relevant, latest, or top.