DEV Community

#evaluation

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Your Agent Has Observability. It Doesn't Have Evals.

Your Agent Has Observability. It Doesn't Have Evals.

Comments
10 min read
LLM Evaluation: How a Benchmark Produces Comparable Numbers

LLM Evaluation: How a Benchmark Produces Comparable Numbers

Comments
4 min read
How LLM Evaluation Actually Works: Inside the Satellite Geo QCM Leaderboard

How LLM Evaluation Actually Works: Inside the Satellite Geo QCM Leaderboard

1
Comments
4 min read
Cosine Gating Won't Save You From Sycophancy: A Self-Refutation Observed by Three Judges

Cosine Gating Won't Save You From Sycophancy: A Self-Refutation Observed by Three Judges

Comments 2
5 min read
An LLM reviewer's "block" is a feature, not a verdict

An LLM reviewer's "block" is a feature, not a verdict

Comments 1
4 min read
My local 7B thinks "kill a Python process" is a violent crime — and my regex beat it

My local 7B thinks "kill a Python process" is a violent crime — and my regex beat it

Comments
4 min read
RAG Evaluation 2026: The Four Core Metrics and How to Read Them Diagnostically

RAG Evaluation 2026: The Four Core Metrics and How to Read Them Diagnostically

Comments
4 min read
Using Execution Traces to Evaluate AI Agent Behavior

Using Execution Traces to Evaluate AI Agent Behavior

2
Comments
5 min read
Cheap LLM code review is fine until it hits an authorization bug

Cheap LLM code review is fine until it hits an authorization bug

Comments
2 min read
7 Agent Eval Mistakes That Cost Me Weeks (And the One-Line Fixes That Ended Them)

7 Agent Eval Mistakes That Cost Me Weeks (And the One-Line Fixes That Ended Them)

24
Comments 5
5 min read
Ranking Isn't Judging: What the Jev Wave Still Owes Retrieval

Ranking Isn't Judging: What the Jev Wave Still Owes Retrieval

1
Comments 1
4 min read
El modelo encontró evidencia relevante y aun así falló: por qué “alucinación” se me quedó corta

El modelo encontró evidencia relevante y aun así falló: por qué “alucinación” se me quedó corta

Comments
4 min read
The sleep loop is the tell: agents that pay per action optimize to do nothing

The sleep loop is the tell: agents that pay per action optimize to do nothing

Comments 1
2 min read
The model did the reverse-engineering. The validator was the hard part.

The model did the reverse-engineering. The validator was the hard part.

Comments 1
2 min read
Judging AI hackathon projects: what to check when every team says 'we used AI'

Judging AI hackathon projects: what to check when every team says 'we used AI'

Comments
3 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.