DEV Community

#localllm

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
Temperature 0 is not reproducible. I measured 30 percent of my output changing between identical runs.

Temperature 0 is not reproducible. I measured 30 percent of my output changing between identical runs.

1
Comments 1
3 min read
Our 4B beat Claude Opus on a 440K-token corpus. Then it came last on the public benchmark.

Our 4B beat Claude Opus on a 440K-token corpus. Then it came last on the public benchmark.

2
Comments
4 min read
I told the model to separate fields with <TAB>. It did exactly that, and I lost 79 percent of my data.

I told the model to separate fields with <TAB>. It did exactly that, and I lost 79 percent of my data.

1
Comments
3 min read
A 4B on a 6GB laptop matched frontier-model accuracy on aggregation — except when the answer is a number

A 4B on a 6GB laptop matched frontier-model accuracy on aggregation — except when the answer is a number

1
Comments
4 min read
Your agent truncates the corpus and answers anyway. Two harnesses, and a router that picks between them.

Your agent truncates the corpus and answers anyway. Two harnesses, and a router that picks between them.

1
Comments
5 min read
What really fits in 8GB VRAM

What really fits in 8GB VRAM

Comments
7 min read
Moving Scheduled LLM Curation from Cloud APIs to Local Models

Moving Scheduled LLM Curation from Cloud APIs to Local Models

Comments
9 min read
Nine ways to talk to a local model

Nine ways to talk to a local model

Comments
9 min read
A 4B model on a 6GB laptop beat Claude Opus on our 440K-token corpus. The fix was giving the model less to do.

A 4B model on a 6GB laptop beat Claude Opus on our 440K-token corpus. The fix was giving the model less to do.

1
Comments 1
4 min read
[Day 20] Local AI vs cloud AI: one cat photo, 10 video models

[Day 20] Local AI vs cloud AI: one cat photo, 10 video models

Comments
3 min read
I A/B tested my own system prompt: 24 generations, one clear win, one rule that did nothing | Flash Onyx 2.2

I A/B tested my own system prompt: 24 generations, one clear win, one rule that did nothing | Flash Onyx 2.2

Comments
8 min read
Flash Onyx 2.2: teaching a local model law and game feel

Flash Onyx 2.2: teaching a local model law and game feel

1
Comments
4 min read
Running an LLM agent entirely in your browser

Running an LLM agent entirely in your browser

Comments
5 min read
Cloud bills kept climbing from 24/7 AI development — I moved the decisions and the implementation to my own local LLMs and cut the cost

Cloud bills kept climbing from 24/7 AI development — I moved the decisions and the implementation to my own local LLMs and cut the cost

Comments
10 min read
Moving Half of Our AI Development to Local LLMs — by Splitting Work by Role, Not by Picking the Biggest Model

Moving Half of Our AI Development to Local LLMs — by Splitting Work by Role, Not by Picking the Biggest Model

Comments 2
3 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.