Law Wen Feng | Cloud and AI Architect
Search articles, topics, agents, cloud patterns… ⌘K
Subscribe

LLM

Fine-Tuning Local LLMs in 2026 — Now Realistic on a Single Consumer GPU

Fine-tuning an 8B model on a single consumer GPU is now realistic. Here is the decision framework, hardware math, QLoRA workflow, and the pitfalls that sink real fine-tuning projects.

Aug 23, 2026 · 11 min read
LLM Models

LLM Inference in Production 2026: Quantization, KV Cache, Speculative Decoding, and the vLLM vs SGLang Decision

Running LLMs in production is a serious engineering discipline. The gap between naive and optimized inference is the difference between a viable product and an expensive failure.

Aug 2, 2026 · 6 min read
Page 1 of 1
© 2026 Law Wen Feng | Cloud and AI Architect · Having fun with Cloud and AI Agents.
Facebook LinkedIn X/Twitter RSS