Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
benchmark
Follow
Hide
Posts
Left menu
ð
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
I Tested Q4_K_M vs MXFP4 on the Same Laptop â The Supposedly-Faster New Format Lost
Pitambar Mahato
Pitambar Mahato
Pitambar Mahato
Follow
Sep 5
I Tested Q4_K_M vs MXFP4 on the Same Laptop â The Supposedly-Faster New Format Lost
#
llm
#
benchmark
#
performance
#
ai
Comments
Add Comment
6 min read
āđāļĄāļ·āđāļ Benchmark āđāļāļŦāļāļāļļāļ, SWE-Bench ProMax āļāļąāļāļāļ°āđāļāļāļāļĢāļīāļāļāļĩāđāđāļĄāđāļāļĨāđāļāđāļāļŠāļļāļāļāļģāđāļāđāđāļāđ 41.2%
Nokka
Nokka
Nokka
Follow
Sep 5
āđāļĄāļ·āđāļ Benchmark āđāļāļŦāļāļāļļāļ, SWE-Bench ProMax āļāļąāļāļāļ°āđāļāļāļāļĢāļīāļāļāļĩāđāđāļĄāđāļāļĨāđāļāđāļāļŠāļļāļāļāļģāđāļāđāđāļāđ 41.2%
#
ai
#
benchmark
#
programming
#
machinelearning
Comments
Add Comment
2 min read
āļāđāļāļāļ§āđāļēāļ 0.3% āđāļāđāļĢāļēāļāļēāļāđāļēāļ 2 āđāļāđāļē, āļāđāļēāļāļāļēāļĢāļēāļ Terminal-Bench 4.0 āđāļŦāđāđāļāđāļ
Nokka
Nokka
Nokka
Follow
Sep 5
āļāđāļāļāļ§āđāļēāļ 0.3% āđāļāđāļĢāļēāļāļēāļāđāļēāļ 2 āđāļāđāļē, āļāđāļēāļāļāļēāļĢāļēāļ Terminal-Bench 4.0 āđāļŦāđāđāļāđāļ
#
ai
#
benchmark
#
llm
#
programming
Comments
Add Comment
2 min read
Benchmarking Real-Time Voice AI APIs: Cartesia vs Deepgram vs ElevenLabs (2026)
mrzitoun
mrzitoun
mrzitoun
Follow
Sep 3
Benchmarking Real-Time Voice AI APIs: Cartesia vs Deepgram vs ElevenLabs (2026)
#
ai
#
webdev
#
voice
#
benchmark
Comments
Add Comment
1 min read
I built a benchmark for AI companion apps because every "best AI girlfriend" list is affiliate spam
Maya Stone
Maya Stone
Maya Stone
Follow
Sep 3
I built a benchmark for AI companion apps because every "best AI girlfriend" list is affiliate spam
#
ai
#
llm
#
benchmark
Comments
1
 comment
5 min read
A Benchmark Is Only as Honest as Its Harness
Avery Wang
Avery Wang
Avery Wang
Follow
Sep 2
A Benchmark Is Only as Honest as Its Harness
#
ai
#
benchmark
#
llm
#
testing
Comments
Add Comment
4 min read
Extraction APIs invented values for 17% of the fields that aren't in the document. Ours included.
Velrim
Velrim
Velrim
Follow
Sep 2
Extraction APIs invented values for 17% of the fields that aren't in the document. Ours included.
#
ai
#
machinelearning
#
benchmark
#
llm
Comments
Add Comment
12 min read
A Small Transformer Trained in 1.5 Hours Beat Many LLMs on ARC
Anaz S. Aji
Anaz S. Aji
Anaz S. Aji
Follow
for
Codecora Dev
Sep 2
A Small Transformer Trained in 1.5 Hours Beat Many LLMs on ARC
#
machinelearning
#
vectorsearch
#
quantization
#
benchmark
Comments
Add Comment
5 min read
I Ran 3 Open-Weight LLMs Head-to-Head on a 24GB Mac â One Was 3x Faster
Pitambar Mahato
Pitambar Mahato
Pitambar Mahato
Follow
Sep 4
I Ran 3 Open-Weight LLMs Head-to-Head on a 24GB Mac â One Was 3x Faster
#
llm
#
benchmark
#
opensource
#
ai
2
 reactions
Comments
4
 comments
5 min read
Benchmark a Free AI Coding Tier on a Cold Server
Avery Wang
Avery Wang
Avery Wang
Follow
Aug 30
Benchmark a Free AI Coding Tier on a Cold Server
#
ai
#
benchmark
#
opensource
#
testing
Comments
Add Comment
4 min read
AxonASP vs. Native IIS ASP: Performance Benchmarks and Engine Architecture
Lucas GuimarÃĢes
Lucas GuimarÃĢes
Lucas GuimarÃĢes
Follow
Aug 29
AxonASP vs. Native IIS ASP: Performance Benchmarks and Engine Architecture
#
iis
#
asp
#
vbscript
#
benchmark
Comments
Add Comment
2 min read
We ran 160 agent tasks across two frameworks. The frameworks tied. Then we changed the model.
benchclawio
benchclawio
benchclawio
Follow
for
benchclaw
Aug 29
We ran 160 agent tasks across two frameworks. The frameworks tied. Then we changed the model.
#
python
#
ai
#
testing
#
benchmark
Comments
Add Comment
4 min read
A LongMemEval-S number you can reproduce
Przemek Marzec
Przemek Marzec
Przemek Marzec
Follow
for
Sovantica
Aug 27
A LongMemEval-S number you can reproduce
#
ai
#
agents
#
benchmark
#
python
Comments
4
 comments
6 min read
I Tested GLM-5.3-Flash and Qwen3.8-Flash on 24 Real Tasks
li wujie
li wujie
li wujie
Follow
Aug 27
I Tested GLM-5.3-Flash and Qwen3.8-Flash on 24 Real Tasks
#
ai
#
llm
#
machinelearning
#
benchmark
Comments
Add Comment
6 min read
Twenty Prompts, One Compiler: A Free AI Server's Failure Matrix
Morgan Ma
Morgan Ma
Morgan Ma
Follow
Aug 22
Twenty Prompts, One Compiler: A Free AI Server's Failure Matrix
#
ai
#
cpp
#
testing
#
benchmark
Comments
Add Comment
5 min read
ð
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account