DEV Community

#evaluation

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
How to Evaluate AI Agents

How to Evaluate AI Agents

2
Comments
7 min read
JuryTrace: make agent-judge failures inspectable

JuryTrace: make agent-judge failures inspectable

Comments 1
4 min read
My board never scored an outage as a regression. My evidence couldn't prove it.

My board never scored an outage as a regression. My evidence couldn't prove it.

Comments
5 min read
We spent two days bisecting a prompt change. The regression was noise.

We spent two days bisecting a prompt change. The regression was noise.

Comments
1 min read
Production-Ready Multi-Turn Evaluation

Production-Ready Multi-Turn Evaluation

Comments
7 min read
A Free Server Is Enough to Test a New Model Before You Trust It

A Free Server Is Enough to Test a New Model Before You Trust It

Comments
3 min read
Writing the code is no longer the bottleneck

Writing the code is no longer the bottleneck

Comments
3 min read
Why AI Benchmarks Mean Less Than You Think

Why AI Benchmarks Mean Less Than You Think

Comments
6 min read
One Quality Score Is a Lie: Split Your RAG Judge Into Retrieval, Groundedness, and Relevance

One Quality Score Is a Lie: Split Your RAG Judge Into Retrieval, Groundedness, and Relevance

1
Comments 1
5 min read
We Almost Deployed a Temporal Knowledge Graph. The Eval Said No.

We Almost Deployed a Temporal Knowledge Graph. The Eval Said No.

Comments
7 min read
Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations

Choosing the Right LLM-as-a-Judge: A Practical Guide with Model Recommendations

Comments
7 min read
How EvalPort's Grader System Works: 11 Types for LLM Evaluation

How EvalPort's Grader System Works: 11 Types for LLM Evaluation

Comments
2 min read
Measure the Judge Before You Trust It: Self-Consistency Comes Before Human Agreement

Measure the Judge Before You Trust It: Self-Consistency Comes Before Human Agreement

1
Comments 2
6 min read
RAG Beyond the Demo: Pipeline, Citations, Evaluation, and When Not to Bother

RAG Beyond the Demo: Pipeline, Citations, Evaluation, and When Not to Bother

Comments 2
7 min read
OpenEval: Why LLM Evaluation Needs a Standard Format

OpenEval: Why LLM Evaluation Needs a Standard Format

Comments
1 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.