DEV Community

Kaggle Benchmarking Challenge

This is the official tag for submissions and announcements related to the Kaggle Benchmarking Challenge.

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
10 Real Airflow Incidents, 5 AI Models — Which One Handled Production Best?

Kaggle Benchmarking Challenge Submission

10 Real Airflow Incidents, 5 AI Models — Which One Handled Production Best?

Comments 1
4 min read
Playoff Probability Calibration — LLMs vs. a Real Monte Carlo Model

Playoff Probability Calibration — LLMs vs. a Real Monte Carlo Model

Comments 1
4 min read
When a Failed Request Must Stay Failed: Reservation Replay

Kaggle Benchmarking Challenge Submission

When a Failed Request Must Stay Failed: Reservation Replay

1
Comments 2
3 min read
Do LLMs Actually Fix Tricky React Hooks, or Do They Just Cheat?

Do LLMs Actually Fix Tricky React Hooks, or Do They Just Cheat?

1
Comments
3 min read
A Paused AI Workflow: Retry, Resume, or Keep Holding?

Kaggle Benchmarking Challenge Submission

A Paused AI Workflow: Retry, Resume, or Keep Holding?

Comments
4 min read
An AI Correctly Ignored a Forum Rumor. I Removed One Label and It Paid Out $150.

Kaggle Benchmarking Challenge Submission

An AI Correctly Ignored a Forum Rumor. I Removed One Label and It Paid Out $150.

5
Comments 1
6 min read
MY Edit or Abstain: AI Code-Repair Decision Benchmark

MY Edit or Abstain: AI Code-Repair Decision Benchmark

Comments
3 min read
I put one wrong test in the file. Most models sided with the test.

Kaggle Benchmarking Challenge Submission

I put one wrong test in the file. Most models sided with the test.

Comments 1
6 min read
Only What I Asked

Only What I Asked

Comments
3 min read
Hook constraint benchmark kaggle-challenge DONE !!!

Kaggle Benchmarking Challenge Submission

Hook constraint benchmark kaggle-challenge DONE !!!

Comments
3 min read
The AI Was Right. The Answer Was Still Wrong.

The AI Was Right. The Answer Was Still Wrong.

5
Comments 1
3 min read
UMKM-Bench: I tested 8 LLMs on how Indonesians really text online shops

Kaggle Benchmarking Challenge Submission

UMKM-Bench: I tested 8 LLMs on how Indonesians really text online shops

Comments
8 min read
Small AI Models Can See the Trap. They Fall In Anyway

Kaggle Benchmarking Challenge Submission

Small AI Models Can See the Trap. They Fall In Anyway

1
Comments
5 min read
Can an AI Agent Know When Not to Act? A Fail-Closed Reliability Benchmark Across Six Models

Kaggle Benchmarking Challenge Submission

Can an AI Agent Know When Not to Act? A Fail-Closed Reliability Benchmark Across Six Models

Comments
4 min read
Do LLMs Actually Check Their Tools? I Built a Benchmark That Lies to Them

Kaggle Benchmarking Challenge Submission

Do LLMs Actually Check Their Tools? I Built a Benchmark That Lies to Them

2
Comments 5
8 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.