DEV Community

Riley Wang profile picture

Riley Wang

Software engineer by day, open source contributor by night.

Location Singapore Joined Joined on 
Orphan Tool Calls Look Like Hangs. Pair Every Result.

Orphan Tool Calls Look Like Hangs. Pair Every Result.

Comments
7 min read
Nested Agent Failures Hide in Flat Logs. Build a Span Tree.

Nested Agent Failures Hide in Flat Logs. Build a Span Tree.

Comments
7 min read
Concurrent Tool Calls Scramble Logs. Diff the Span DAG Instead.

Concurrent Tool Calls Scramble Logs. Diff the Span DAG Instead.

Comments
7 min read
Your Agent Is Leaking Tokens. Build a Decision Ledger on a Free Server.

Your Agent Is Leaking Tokens. Build a Decision Ledger on a Free Server.

Comments
4 min read
When the Token Well Runs Dry: A Degradation State Machine for LLM Services

When the Token Well Runs Dry: A Degradation State Machine for LLM Services

Comments
4 min read
How I Built a Self-Auditing Agent Loop on a Free Server

How I Built a Self-Auditing Agent Loop on a Free Server

Comments
5 min read
Log Where Each Tool Argument Came From

Log Where Each Tool Argument Came From

Comments
7 min read
Why Your Agent Fails Only in Production: A Trace-and-Diff Loop on a Free Server

Why Your Agent Fails Only in Production: A Trace-and-Diff Loop on a Free Server

Comments
4 min read
The Agent's Key Ring: Scoping Tool Permissions with a Deny-by-Default Wrapper

The Agent's Key Ring: Scoping Tool Permissions with a Deny-by-Default Wrapper

Comments
4 min read
Diff Every Tool Call: Replaying Agent Runs from a JSONL Trace

Diff Every Tool Call: Replaying Agent Runs from a JSONL Trace

5
Comments 3
4 min read
Your Agent's Context Is Rotting. Here's How I Traced the Decay.

Your Agent's Context Is Rotting. Here's How I Traced the Decay.

Comments 1
4 min read
Trace Agent Tool Calls on a Free Server: A 10M-Token Debug Loop

Trace Agent Tool Calls on a Free Server: A 10M-Token Debug Loop

1
Comments
4 min read
Trace First, Blame Later: A Debug Loop for Failing Agent Runs

Trace First, Blame Later: A Debug Loop for Failing Agent Runs

1
Comments
5 min read
Agent Runs Fail Quietly. Here's the Trace Harness I Use to Find Out Why.

Agent Runs Fail Quietly. Here's the Trace Harness I Use to Find Out Why.

Comments
4 min read
Build a Token-Budgeted LLM Service on a Free Server: A Step-by-Step Tutorial

Build a Token-Budgeted LLM Service on a Free Server: A Step-by-Step Tutorial

Comments
5 min read
Token Forensics: Auditing LLM Usage on a Free Server

Token Forensics: Auditing LLM Usage on a Free Server

Comments
5 min read
MiniMax H3 Is Making Noise. My First Check Isn’t the Leaderboard.

MiniMax H3 Is Making Noise. My First Check Isn’t the Leaderboard.

Comments
3 min read
A New Cheap Model Dropped This Week. Here's the Harness I Run Before I Switch Anything

A New Cheap Model Dropped This Week. Here's the Harness I Run Before I Switch Anything

Comments
5 min read
I Stopped Trusting My Agent's Boundaries Until I Could Break Them in a Throwaway Sandbox

I Stopped Trusting My Agent's Boundaries Until I Could Break Them in a Throwaway Sandbox

Comments
5 min read
Stop Benchmarking Coding Models by Vibes: A Repeatable 20-Task Harness You Can Run Tonight

Stop Benchmarking Coding Models by Vibes: A Repeatable 20-Task Harness You Can Run Tonight

2
Comments
5 min read
loading...