
Episode #28
You've Bought AI hardware, now what?
READ THE FULL EPISODE PAGE https://devmesh.tech/podcast/youve-bought-ai-hardware-now-what Buying AI hardware is only the beginning. In Episode 28 of System Prompt, Peter and Val follow up on Before You Buy AI Hardware, Watch This by moving from planning into the actual hosting layer. Using GLM-5.3-Flash across two NVIDIA DGX Sparks with vLLM, they break down what the serving command is doing, which settings actually matter, and why hosting a model is very different from simply loading one. The conversation covers tensor parallelism, distributed execution, RoCE and NCCL networking, Docker, model mounts, KV cache, context length, concurrency, Mixture-of-Experts execution, chat templates, reasoning and tool-call parsers, DFlash speculative decoding, prefill, decode, and live serving metrics. Then they start the real cluster, bring the worker online before the head node, watch the model load across both Sparks, and run GLM-5.3-Flash through a coding workload. The live metrics show speculative decoding acceptance changing with the task, with generation throughput moving from the teens into the 30s and 40s as the workload becomes more predictable. • What happens after you buy AI hardware • Hosting GLM-5.3-Flash across two DGX Sparks • vLLM and distributed inference • Tensor parallelism across multiple devices • RoCE, NCCL, and cluster communication • Docker, model mounts, and runtime configuration • Context windows and KV cache • FP8 KV cache vs. FP16 • Concurrency and max active sequences • Mixture-of-Experts execution • Chat templates, reasoning parsers, and tool-call parsers • DFlash speculative decoding • Prefill, decode, and generation throughput • Why real workloads have to be tested QUICK CORRECTION At one point I say the ConnectX-7 link is 200 gigabytes per second. I meant 200 gigabits per second, or 200 Gb/s. Yes, I know the difference. It was a live recording, I misspoke, and I am correcting it here. KEY TAKEAWAYS RUNNING A MODEL IS NOT THE SAME AS HOSTING ONE Loading the weights is step one. Hosting adds networking, scheduling, memory management, concurrency, cache behavior, runtime compatibility, and observability. THE MODEL IS ONLY ONE PART OF THE STACK Runtime, quantization, KV cache, speculative decoding, context limits, scheduling, and hardware topology all affect system behavior. YOU DO NOT NEED TO ENGINEER EVERY LAYER YOURSELF Start with a known-good path, understand the major knobs, test it on your hardware, and go deeper only where the workload proves you need to. SPECULATIVE DECODING ONLY HELPS WHEN THE DRAFT IS RIGHT When DFlash acceptance is low, speculative work can become overhead. When the workload becomes more predictable, acceptance and generation throughput can rise significantly. THE WORKLOAD DETERMINES THE CONFIGURATION There is no universal best setup. Context, concurrency, precision, cache size, and speculative decoding all need to be tested against the actual task. CHAPTERS 00:00 Intro and the Follow-Up 03:23 Two DGX Sparks, vLLM, and GLM-5.3-Flash 07:00 Breaking Down the Hosting Command 11:00 Model Mounts, DFlash, and Runtime Patches 14:12 NCCL, RoCE, and Distributed Networking 17:30 API Behavior, Tool Calls, Reasoning, and Chat Templates 21:40 The vLLM Knobs You Actually Tune 25:45 FP8 KV Cache and Memory Tradeoffs 27:30 DFlash and Speculative Decoding 31:38 Starting the Two-Node Cluster 41:51 Running GLM-5.3-Flash 45:56 Reading the Live Metrics 50:17 Why You Have to Test 54:22 Coding Workloads and Higher Draft Acceptance 59:37 What This Means for Businesses 1:04:01 Closing Thoughts

