Skip to content
Navigation menu
Search
Powered by Algolia
Search
Log in
Create account
DEV Community
Close
#
quantization
Follow
Hide
Posts
Left menu
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
Right menu
A Small Transformer Trained in 1.5 Hours Beat Many LLMs on ARC
Anaz S. Aji
Anaz S. Aji
Anaz S. Aji
Follow
for
Codecora Dev
Sep 2
A Small Transformer Trained in 1.5 Hours Beat Many LLMs on ARC
#
machinelearning
#
vectorsearch
#
quantization
#
benchmark
Comments
Add Comment
5 min read
A Better FP4 Gradient Quantizer That Training Couldn't Notice
Seth Wheeler
Seth Wheeler
Seth Wheeler
Follow
Aug 25
A Better FP4 Gradient Quantizer That Training Couldn't Notice
#
llm
#
measurement
#
quantization
#
training
Comments
Add Comment
7 min read
Error Feedback, Gradient Compression, and Why Adam Breaks It
Seth Wheeler
Seth Wheeler
Seth Wheeler
Follow
Aug 21
Error Feedback, Gradient Compression, and Why Adam Breaks It
#
llm
#
measurement
#
quantization
#
training
5
 reactions
Comments
1
 comment
8 min read
KV Cache INT4 Quantization for 1M+ Token Context Windows
Aomi Qaza
Aomi Qaza
Aomi Qaza
Follow
Aug 15
KV Cache INT4 Quantization for 1M+ Token Context Windows
#
aiengineering
#
quantization
Comments
Add Comment
3 min read
Comparing INT4 and NVFP4 Palettes on Real Gradient Tensors
Seth Wheeler
Seth Wheeler
Seth Wheeler
Follow
Aug 22
Comparing INT4 and NVFP4 Palettes on Real Gradient Tensors
#
llm
#
quantization
#
training
#
measurement
1
 reaction
Comments
2
 comments
7 min read
Deploying a QAT Checkpoint Your Serving Stack Can't Load: Gemma 4 E2B in Pure JAX on One TPU
xbill
xbill
xbill
Follow
for
Google Developer Experts
Aug 19
Deploying a QAT Checkpoint Your Serving Stack Can't Load: Gemma 4 E2B in Pure JAX on One TPU
#
tpu
#
jax
#
llm
#
quantization
8
 reactions
Comments
Add Comment
10 min read
Unsloth Releases Qwen3.6-27B-NVFP4: Enhanced Throughput and Agentic Coding for Developers
Pneumetron
Pneumetron
Pneumetron
Follow
Jul 17
Unsloth Releases Qwen3.6-27B-NVFP4: Enhanced Throughput and Agentic Coding for Developers
#
aiml
#
largelanguagemodels
#
quantization
#
qwen
Comments
Add Comment
3 min read
Bonsai-27B: A 1-Bit LLM for On-Device Inference with Llama.cpp and MLX
Pneumetron
Pneumetron
Pneumetron
Follow
Jul 15
Bonsai-27B: A 1-Bit LLM for On-Device Inference with Llama.cpp and MLX
#
llm
#
quantization
#
1bit
#
gguf
Comments
Add Comment
3 min read
Quantizing MedGemma to INT4 (GPTQ/W4A16): Everything That Broke Along the Way
JoTeq the First
JoTeq the First
JoTeq the First
Follow
Jul 14
Quantizing MedGemma to INT4 (GPTQ/W4A16): Everything That Broke Along the Way
#
machinelearning
#
llm
#
quantization
#
opensource
Comments
Add Comment
6 min read
LLM Inference Optimization: From Quantization to Speculative Decoding
Aarush Karak
Aarush Karak
Aarush Karak
Follow
Aug 13
LLM Inference Optimization: From Quantization to Speculative Decoding
#
llm
#
inference
#
quantization
#
optimization
1
 reaction
Comments
1
 comment
2 min read
Self-Hosting a Model Means Self-Hosting Its Evaluation Too
AI Explore
AI Explore
AI Explore
Follow
Jul 7
Self-Hosting a Model Means Self-Hosting Its Evaluation Too
#
ai
#
llm
#
quantization
#
mlops
1
 reaction
Comments
2
 comments
5 min read
Gemma 4 QAT on a 1080 Ti: What 'Quantization-Aware' Actually Buys — and Fitting the 12B on 8 GB at 16k
byeongsoo kang
byeongsoo kang
byeongsoo kang
Follow
Jun 11
Gemma 4 QAT on a 1080 Ti: What 'Quantization-Aware' Actually Buys — and Fitting the 12B on 8 GB at 16k
#
llm
#
machinelearning
#
gemma
#
quantization
Comments
Add Comment
5 min read
Quantization formats compared: GGUF vs GPTQ vs AWQ vs NF4
Tech_Nuggets
Tech_Nuggets
Tech_Nuggets
Follow
Jun 11
Quantization formats compared: GGUF vs GPTQ vs AWQ vs NF4
#
llm
#
quantization
#
mlops
#
tutorial
Comments
Add Comment
7 min read
LLM Quantization Levels Compared: Q4_K_M vs Q8_0 vs FP16 [2026]
Kunal
Kunal
Kunal
Follow
Jul 6
LLM Quantization Levels Compared: Q4_K_M vs Q8_0 vs FP16 [2026]
#
localllm
#
quantization
#
gguf
#
ollama
Comments
1
 comment
15 min read
How to Pick a GGUF Quant Level for Your VRAM Budget
Patrick Hughes
Patrick Hughes
Patrick Hughes
Follow
Jun 11
How to Pick a GGUF Quant Level for Your VRAM Budget
#
localllm
#
gguf
#
quantization
#
gpu
Comments
Add Comment
4 min read
đź‘‹
Sign in
for the ability to sort posts by
relevant
,
latest
, or
top
.
We're a place where coders share, stay up-to-date and grow their careers.
Log in
Create account