Law Wen Feng | Cloud and AI Architect
Search articles, topics, agents, cloud patterns… ⌘K
Subscribe

LLM

LLM Inference Optimization 2026: The Shift from Software Tweaks to Hardware-Software Co-Design

Why LLM inference optimization moved from software tweaks to hardware-software co-design, and what it means for Azure-first enterprise teams in 2026.

Aug 19, 2026 · 13 min read
LLM Models

LLM Inference in Production 2026: Quantization, KV Cache, Speculative Decoding, and the vLLM vs SGLang Decision

Running LLMs in production is a serious engineering discipline. The gap between naive and optimized inference is the difference between a viable product and an expensive failure.

Aug 2, 2026 · 6 min read
Page 1 of 1
© 2026 Law Wen Feng | Cloud and AI Architect · Having fun with Cloud and AI Agents.
Facebook LinkedIn X/Twitter RSS