DEV Community

Casey Zhang profile picture

Casey Zhang

Building things with Python, Go, and JavaScript. Love automating everything.

Location Singapore Joined Joined on 
Don't Ship the Pass Rate. Ship the Failure Mix.

Don't Ship the Pass Rate. Ship the Failure Mix.

Comments
7 min read
Freeze the Retry Budget Before You Trust an Agent Pass Rate

Freeze the Retry Budget Before You Trust an Agent Pass Rate

Comments
8 min read
Build an LLM Cache Proxy in 40 Lines to Cut Token Costs 83%

Build an LLM Cache Proxy in 40 Lines to Cut Token Costs 83%

Comments
6 min read
Your Coding Agent's Pass Rate Isn't a Benchmark Until the Oracle Is Frozen

Your Coding Agent's Pass Rate Isn't a Benchmark Until the Oracle Is Frozen

Comments
7 min read
Benchmarking an AI Coding Tool? Start With the Data You Actually Trust

Benchmarking an AI Coding Tool? Start With the Data You Actually Trust

Comments
5 min read
Your Prompt A/B Test Is Probably Random Noise: A Paired-Test Harness for Free-Tier LLMs

Your Prompt A/B Test Is Probably Random Noise: A Paired-Test Harness for Free-Tier LLMs

Comments
5 min read
Benchmarking AI Code Generators: A Reproducible Method You Can Run on a Free Server

Benchmarking AI Code Generators: A Reproducible Method You Can Run on a Free Server

Comments
5 min read
Cold Starts Will Eat Your Free LLM Tier: A Reproducible Benchmark

Cold Starts Will Eat Your Free LLM Tier: A Reproducible Benchmark

Comments
4 min read
The Token Allowance Trap: A Task-Price Benchmark Before You Build on Free LLM APIs

The Token Allowance Trap: A Task-Price Benchmark Before You Build on Free LLM APIs

Comments
4 min read
Measure Your AI Reviewer Before You Trust It: A Seeded-Defect Benchmark

Measure Your AI Reviewer Before You Trust It: A Seeded-Defect Benchmark

Comments
5 min read
Benchmarking LLM Agents Without the Marketing Math: Dataset, Metrics, and Controls

Benchmarking LLM Agents Without the Marketing Math: Dataset, Metrics, and Controls

Comments
5 min read
From git log to Release Notes: A Free-Tier Agent Case Study

From git log to Release Notes: A Free-Tier Agent Case Study

Comments
5 min read
Crash-Proof Batch LLM Processing: A SQLite Job Queue on a Free Server

Crash-Proof Batch LLM Processing: A SQLite Job Queue on a Free Server

Comments 1
5 min read
Building a Free-Tier GitHub Issue Triage Agent: A Case Study

Building a Free-Tier GitHub Issue Triage Agent: A Case Study

Comments
5 min read
Cheap Model Hype Is Not a Benchmark: Run a 30-Minute Agent Gate Before You Swap

Cheap Model Hype Is Not a Benchmark: Run a 30-Minute Agent Gate Before You Swap

Comments
4 min read
Cheap Model Hype Is Not a Benchmark: Run a 30-Minute Agent Gate Before You Swap

Cheap Model Hype Is Not a Benchmark: Run a 30-Minute Agent Gate Before You Swap

Comments
4 min read
Tool-Permission Gate: The Only MiniMax H3 Test That Matters

Tool-Permission Gate: The Only MiniMax H3 Test That Matters

Comments
5 min read
Pin Your Agent's Behavior: Snapshot Tests for Tool-Call Sequences When You Swap Models

Pin Your Agent's Behavior: Snapshot Tests for Tool-Call Sequences When You Swap Models

Comments
5 min read
AI Agent Boundary Testing: Fake-Tool Harness on a Free Server

AI Agent Boundary Testing: Fake-Tool Harness on a Free Server

Comments
6 min read
I Gave My Coding Agent Fake Tools on Purpose: A Boundary-Test Harness You Can Run on a Free Server

I Gave My Coding Agent Fake Tools on Purpose: A Boundary-Test Harness You Can Run on a Free Server

Comments
5 min read
Stop Guessing: A Reproducible Harness for Evaluating Free AI Coding Models on Your Own Repo

Stop Guessing: A Reproducible Harness for Evaluating Free AI Coding Models on Your Own Repo

Comments
5 min read
loading...