Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
← All Trends
Red-Teaming AI Coding Agent Sandboxes
41 posts in this trend in the last 7 days
•
Active about 5 hours ago
A Preflight Harness for Tool-Using Coding Agents
Sam Sun
Sam Sun
Sam Sun
Follow
Aug 10
A Preflight Harness for Tool-Using Coding Agents
#
ai
#
security
#
programming
#
productivity
Comments
1
comment
5 min read
A Boundary-Failure Test Plan for Coding Agents You Can Run on Free Model Tiers
Dakota Wu
Dakota Wu
Dakota Wu
Follow
Aug 10
A Boundary-Failure Test Plan for Coding Agents You Can Run on Free Model Tiers
#
security
#
ai
#
agents
#
testing
Comments
Add Comment
5 min read
A Disposable Sandbox Pattern for Testing AI Coding Agents Safely
Taylor Wang
Taylor Wang
Taylor Wang
Follow
Aug 10
A Disposable Sandbox Pattern for Testing AI Coding Agents Safely
#
ai
#
security
#
productivity
#
tutorial
Comments
Add Comment
4 min read
Canary Tests for AI Coding Agents: A Sandbox Harness You Can Run Yourself
Harper Zhu
Harper Zhu
Harper Zhu
Follow
Aug 7
Canary Tests for AI Coding Agents: A Sandbox Harness You Can Run Yourself
#
ai
#
agents
#
security
#
testing
Comments
Add Comment
4 min read
AI Agent Boundary Testing: Fake-Tool Harness on a Free Server
Casey Zhang
Casey Zhang
Casey Zhang
Follow
Aug 7
AI Agent Boundary Testing: Fake-Tool Harness on a Free Server
#
aiagents
#
security
#
promptinjection
#
python
Comments
Add Comment
6 min read
A Reproducible Sandbox Loop for AI-Generated Code: Generate, Isolate, Assert
Charlie Zhu
Charlie Zhu
Charlie Zhu
Follow
Aug 10
A Reproducible Sandbox Loop for AI-Generated Code: Generate, Isolate, Assert
#
ai
#
security
#
testing
#
tutorial
Comments
Add Comment
4 min read
Canaries, Not Faith: Auditing Where Your Coding Agent Actually Writes
Sam Yang
Sam Yang
Sam Yang
Follow
Aug 7
Canaries, Not Faith: Auditing Where Your Coding Agent Actually Writes
#
ai
#
security
#
agents
#
python
Comments
Add Comment
6 min read
Don't Trust the Transcript: A Pytest Harness That Audits What Your AI Coding Agent Actually Did
Emery Chen
Emery Chen
Emery Chen
Follow
Aug 10
Don't Trust the Transcript: A Pytest Harness That Audits What Your AI Coding Agent Actually Did
#
ai
#
security
#
testing
#
agents
Comments
Add Comment
6 min read
Your Agent's Permission Slip Belongs in Version Control, Not in a Prompt
Jordan Liu
Jordan Liu
Jordan Liu
Follow
Aug 10
Your Agent's Permission Slip Belongs in Version Control, Not in a Prompt
#
security
#
ai
#
agents
#
testing
Comments
Add Comment
7 min read
Red-Teaming AI Coding Agents Without a Budget: A Boundary Test Suite on Free Models
Harper Zhu
Harper Zhu
Harper Zhu
Follow
Aug 10
Red-Teaming AI Coding Agents Without a Budget: A Boundary Test Suite on Free Models
#
ai
#
agents
#
security
#
testing
Comments
Add Comment
6 min read
Your Coding Agent Has More Access Than You Think. Here's the Audit.
Build Loops
Build Loops
Build Loops
Follow
Aug 10
Your Coding Agent Has More Access Than You Think. Here's the Audit.
#
ai
#
agents
#
security
#
programming
Comments
1
comment
7 min read
I Stopped Running Agent-Generated Code on My Laptop: A 20-Minute Disposable Sandbox
Sam Rivera
Sam Rivera
Sam Rivera
Follow
Aug 6
I Stopped Running Agent-Generated Code on My Laptop: A 20-Minute Disposable Sandbox
#
ai
#
agents
#
security
#
tutorial
Comments
1
comment
5 min read
Your System Prompt Is Not a Security Boundary: A Hands-On Probe for Agent Tool Calls
Casey Chen
Casey Chen
Casey Chen
Follow
Aug 10
Your System Prompt Is Not a Security Boundary: A Hands-On Probe for Agent Tool Calls
#
ai
#
security
#
agents
#
python
Comments
Add Comment
6 min read
Don't Trust-Execute AI-Generated Code: A Sandbox Harness for Evaluating Coding Models Safely
Jordan Li
Jordan Li
Jordan Li
Follow
Aug 7
Don't Trust-Execute AI-Generated Code: A Sandbox Harness for Evaluating Coding Models Safely
#
ai
#
security
#
docker
#
programming
Comments
Add Comment
6 min read
Your AI Coding Assistant Writes Shell Commands. Do You Actually Test Them Before They Run?
Avery Wang
Avery Wang
Avery Wang
Follow
Aug 10
Your AI Coding Assistant Writes Shell Commands. Do You Actually Test Them Before They Run?
#
ai
#
security
#
productivity
#
bash
Comments
Add Comment
5 min read
Sandbox First: A Safer Local Harness for Evaluating Free Coding Models on Your Own Codebase
Avery Lin
Avery Lin
Avery Lin
Follow
Aug 7
Sandbox First: A Safer Local Harness for Evaluating Free Coding Models on Your Own Codebase
#
ai
#
programming
#
security
#
testing
Comments
Add Comment
6 min read
Your Agent Reads Untrusted Text All Day. Here's How I Grade What It Does With It.
Dakota Ma
Dakota Ma
Dakota Ma
Follow
Aug 10
Your Agent Reads Untrusted Text All Day. Here's How I Grade What It Does With It.
#
security
#
ai
#
agents
#
testing
Comments
Add Comment
6 min read
Model Swaps Are Boundary Events: Gate Agent Tool Changes With a Deterministic Replay Lane
Casey Sun
Casey Sun
Casey Sun
Follow
Aug 10
Model Swaps Are Boundary Events: Gate Agent Tool Changes With a Deterministic Replay Lane
#
ai
#
security
#
agents
#
testing
Comments
Add Comment
6 min read
« First
‹ Prev
1
2
3
Next ›
Last »
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account