Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
evaluation
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
Your Agent Has Observability. It Doesn't Have Evals.
Jason Lau
Jason Lau
Jason Lau
Follow
Sep 24
Your Agent Has Observability. It Doesn't Have Evals.
#
agents
#
observability
#
evaluation
#
llm
Comments
Add Comment
10 min read
LLM Evaluation: How a Benchmark Produces Comparable Numbers
RESK
RESK
RESK
Follow
Sep 24
LLM Evaluation: How a Benchmark Produces Comparable Numbers
#
llm
#
evaluation
#
benchmark
#
reinforcementlearnin
Comments
Add Comment
4 min read
How LLM Evaluation Actually Works: Inside the Satellite Geo QCM Leaderboard
RESK
RESK
RESK
Follow
Sep 22
How LLM Evaluation Actually Works: Inside the Satellite Geo QCM Leaderboard
#
llm
#
evaluation
#
benchmark
#
vision
1
 reaction
Comments
Add Comment
4 min read
Cosine Gating Won't Save You From Sycophancy: A Self-Refutation Observed by Three Judges
Modusensus
Modusensus
Modusensus
Follow
Sep 25
Cosine Gating Won't Save You From Sycophancy: A Self-Refutation Observed by Three Judges
#
ai
#
oss
#
evaluation
#
llm
Comments
2
 comments
5 min read
An LLM reviewer's "block" is a feature, not a verdict
Cole Halton
Cole Halton
Cole Halton
Follow
Sep 19
An LLM reviewer's "block" is a feature, not a verdict
#
aicodereview
#
llm
#
evaluation
#
machinelearning
Comments
1
 comment
4 min read
My local 7B thinks "kill a Python process" is a violent crime — and my regex beat it
Amirul Cyber
Amirul Cyber
Amirul Cyber
Follow
Sep 17
My local 7B thinks "kill a Python process" is a violent crime — and my regex beat it
#
ai
#
security
#
evaluation
#
llm
Comments
Add Comment
4 min read
RAG Evaluation 2026: The Four Core Metrics and How to Read Them Diagnostically
saaro
saaro
saaro
Follow
Sep 17
RAG Evaluation 2026: The Four Core Metrics and How to Read Them Diagnostically
#
rag
#
ai
#
evaluation
#
metrics
Comments
Add Comment
4 min read
Using Execution Traces to Evaluate AI Agent Behavior
Quantiles.io
Quantiles.io
Quantiles.io
Follow
Sep 17
Using Execution Traces to Evaluate AI Agent Behavior
#
ai
#
opensource
#
evaluation
2
 reactions
Comments
Add Comment
5 min read
Cheap LLM code review is fine until it hits an authorization bug
Cole Halton
Cole Halton
Cole Halton
Follow
Sep 15
Cheap LLM code review is fine until it hits an authorization bug
#
aicodereview
#
llm
#
security
#
evaluation
Comments
Add Comment
2 min read
7 Agent Eval Mistakes That Cost Me Weeks (And the One-Line Fixes That Ended Them)
Debashish Ghosal
Debashish Ghosal
Debashish Ghosal
Follow
Sep 24
7 Agent Eval Mistakes That Cost Me Weeks (And the One-Line Fixes That Ended Them)
#
ai
#
evaluation
#
llm
#
programming
24
 reactions
Comments
5
 comments
5 min read
Ranking Isn't Judging: What the Jev Wave Still Owes Retrieval
Igor Eduardo
Igor Eduardo
Igor Eduardo
Follow
Sep 24
Ranking Isn't Judging: What the Jev Wave Still Owes Retrieval
#
rag
#
evaluation
#
ai
1
 reaction
Comments
1
 comment
4 min read
El modelo encontró evidencia relevante y aun asà falló: por qué “alucinación” se me quedó corta
Cristian Gormaz
Cristian Gormaz
Cristian Gormaz
Follow
Sep 7
El modelo encontró evidencia relevante y aun asà falló: por qué “alucinación” se me quedó corta
#
ai
#
testing
#
llm
#
evaluation
Comments
Add Comment
4 min read
The sleep loop is the tell: agents that pay per action optimize to do nothing
Cole Halton
Cole Halton
Cole Halton
Follow
Sep 7
The sleep loop is the tell: agents that pay per action optimize to do nothing
#
aiagents
#
evaluation
#
llm
#
benchmarking
Comments
1
 comment
2 min read
The model did the reverse-engineering. The validator was the hard part.
Cole Halton
Cole Halton
Cole Halton
Follow
Sep 7
The model did the reverse-engineering. The validator was the hard part.
#
aicoding
#
agents
#
llm
#
evaluation
Comments
1
 comment
2 min read
Judging AI hackathon projects: what to check when every team says 'we used AI'
PRANJUL RATHOUR
PRANJUL RATHOUR
PRANJUL RATHOUR
Follow
Sep 6
Judging AI hackathon projects: what to check when every team says 'we used AI'
#
hackathonjudging
#
ai
#
evaluation
#
rubric
Comments
Add Comment
3 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account