DEV Community

#gpu

Posts

đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.
From API to GPU, Week 5: Tensors, the Data Structure Behind Every Model

From API to GPU, Week 5: Tensors, the Data Structure Behind Every Model

Comments
12 min read
Mastering Low-Precision AI: FP8 and FP4 Support Across Frameworks in Mid-2026

Mastering Low-Precision AI: FP8 and FP4 Support Across Frameworks in Mid-2026

Comments
2 min read
Rust Portable SIMD Now Runs on the GPU and It Changes Everything About Cross-Platform Parallelism

Rust Portable SIMD Now Runs on the GPU and It Changes Everything About Cross-Platform Parallelism

Comments
4 min read
Understanding GPU Memory: VRAM, Bandwidth, and Why Your Model Won't Fit

Understanding GPU Memory: VRAM, Bandwidth, and Why Your Model Won't Fit

1
Comments
2 min read
GPU_WORKLOAD_MISMATCH Part II: From Detection to Runtime Enforcement for AI Infrastructure

GPU_WORKLOAD_MISMATCH Part II: From Detection to Runtime Enforcement for AI Infrastructure

1
Comments 1
9 min read
Rust SIMD Just Came to the GPU — and It Changes How We Think About Parallel Programming

Rust SIMD Just Came to the GPU — and It Changes How We Think About Parallel Programming

2
Comments
4 min read
How We Cut Inference Cold Starts from Minutes to Seconds

How We Cut Inference Cold Starts from Minutes to Seconds

1
Comments
6 min read
Renting GPUs for AI? Start with VRAM, Not the GPU

Renting GPUs for AI? Start with VRAM, Not the GPU

Comments
2 min read
Why memory bandwidth matters more than TFLOPS for LLM inference

Why memory bandwidth matters more than TFLOPS for LLM inference

Comments
3 min read
Deploying DeepSeek V3 (LLM) Using SGLang

Deploying DeepSeek V3 (LLM) Using SGLang

6
Comments 1
2 min read
One GPU, four ways to share it: ten scenarios, and the headline finding I had to retract

One GPU, four ways to share it: ten scenarios, and the headline finding I had to retract

1
Comments 1
8 min read
KV Cache Quantization: I Stretched Qwen 35B's Context 8 on 12GB VRAM

KV Cache Quantization: I Stretched Qwen 35B's Context 8 on 12GB VRAM

1
Comments 1
3 min read
GPU Monitoring & Metrics for MLOps

GPU Monitoring & Metrics for MLOps

Comments
1 min read
The 60% idle GPU that turned out to be a network policy

The 60% idle GPU that turned out to be a network policy

Comments
3 min read
Accidentally quadratic: buffer copies made MCTS in DeepMind's mctx 3 slower

Accidentally quadratic: buffer copies made MCTS in DeepMind's mctx 3 slower

1
Comments
6 min read
đź‘‹ Sign in for the ability to sort posts by relevant, latest, or top.