[{"data":1,"prerenderedAt":5114},["ShallowReactive",2],{"blog-beating-llamacpp-from-scratch-consumer-amd":3,"blog-beating-llamacpp-from-scratch-consumer-amd-favor-of":5113},{"id":4,"title":5,"body":6,"cover_image":5091,"date":5092,"description":5093,"extension":5094,"meta":5095,"navigation":1913,"path":5096,"seo":5097,"stem":5098,"tags":5099,"type":5111,"__hash__":5112},"blog\u002Fblog\u002Fbeating-llamacpp-from-scratch-consumer-amd.md","Beating llama.cpp from Scratch on Consumer AMD: Building Strata, a Native Rust + HIP Local AI Engine",{"type":7,"value":8,"toc":5055},"minimark",[9,13,31,37,44,56,66,69,90,110,113,262,276,282,285,288,293,296,459,461,465,539,546,880,892,895,903,905,909,919,922,932,939,944,950,977,981,1181,1184,1186,1190,1222,1371,1381,1391,1398,1402,1405,1667,1674,1681,1684,1686,1690,1700,1703,1706,1741,1745,1752,1770,1814,1824,1835,1849,1852,1859,1863,1877,1884,1943,1946,2003,2009,2011,2015,2021,2027,2030,2059,2063,2066,2092,2095,2101,2107,2128,2131,2134,2136,2140,2143,2149,2152,2155,2159,2162,2293,2303,2307,2325,2792,2799,2901,2904,3032,3038,3041,3047,3054,3061,3063,3067,3075,3188,3201,3205,3213,3295,3301,3370,3380,3386,3396,3403,3405,3409,3415,3418,3582,3588,3595,3602,3606,3611,3617,3620,4340,4347,4355,4357,4361,4364,4368,4377,4386,4395,4694,4700,4710,4717,4721,4732,4735,4759,4766,4770,4777,4781,4948,4950,4954,4960,4986,4988,4996,4998,5002,5018,5051],[10,11,5],"h1",{"id":12},"beating-llamacpp-from-scratch-on-consumer-amd-building-strata-a-native-rust-hip-local-ai-engine",[14,15,16],"p",{},[17,18,19,20,30],"em",{},"How Software Architect Mihai Farcas engineered ",[21,22,23],"strong",{},[24,25,29],"a",{"href":26,"rel":27},"https:\u002F\u002Fgithub.com\u002Fmihailtd\u002Fstrata",[28],"nofollow","Strata"," for Qwen 3.5 on an AMD Radeon RX 7900 XTX (gfx1100), eliminated a 3,300-kernel prefill launch storm, fixed an un-flushed 8KB socket stall, and achieved 83.0 tok\u002Fs sustained BF16 decode.",[14,32,33,36],{},[21,34,35],{},"By Mihai Farcas"," — Software Architect & AI Systems Engineer",[14,38,39],{},[40,41],"img",{"alt":42,"src":43},"Strata Architecture: Rust Host Orchestration, Qwen 3.5 Hybrid Topology, and AMD RDNA3 Execution | wide","\u002Fimages\u002Fblog\u002Fstrata-architecture-poster.svg#wide",[14,45,46,47,51,52,55],{},"Can you beat ",[48,49,50],"code",{},"llama.cpp"," and ",[48,53,54],{},"Ollama"," by writing a custom local LLM inference engine from scratch in Rust and AMD HIP on consumer hardware?",[14,57,58,59,61,62,65],{},"The conventional wisdom across the local AI and open-source LLM communities says no. ",[48,60,50],{}," represents thousands of person-years of extreme C++ optimization: hand-tuned AVX-512 and AVX2 vector paths, custom GGML tensor kernels, optimized CUDA\u002FHIP backends, and battle-tested memory allocators. On AMD silicon specifically, the common assumption is even more pessimistic: people call ROCm fragile, rank RDNA3 consumer cards (",[48,63,64],{},"gfx1100",") behind CDNA datacenter accelerators, and expect raw HIP kernels written from scratch to bring driver timeouts and kernel panics.",[14,67,68],{},"We decided to test that assumption directly on bare metal.",[14,70,71,72,77,78,81,82,85,86,89],{},"Over the past two months, as a Software Architect exploring the frontiers of bare-metal local AI and GPU kernel development, I engineered ",[21,73,74],{},[24,75,29],{"href":26,"rel":76},[28]," (internally codenamed ",[48,79,80],{},"runtime-next","): a from-scratch local LLM serving runtime written in safe Rust with custom AMD HIP compute kernels, targeting the ",[21,83,84],{},"Qwen 3.5 hybrid architecture"," running on a single consumer desktop GPU: the ",[21,87,88],{},"AMD Radeon RX 7900 XTX (24GB GDDR6, gfx1100)",".",[14,91,92,93,51,102,109],{},"We benchmarked head-to-head against both ",[21,94,95,97,98,101],{},[48,96,50],{}," (",[48,99,100],{},"llama-server",")",[21,103,104,97,106,101],{},[48,105,54],{},[48,107,108],{},"ollama serve"," under strict, verifiable apples-to-apples conditions: independent HTTP daemons on localhost, streaming Server-Sent Events (SSE) over TCP sockets, unquantized byte-identical bfloat16 parameters, and 100% GPU offload on native ROCm 7.2.",[14,111,112],{},"Here are the headline results on the 4B model:",[114,115,116,149],"table",{},[117,118,119],"thead",{},[120,121,122,127,134,140,146],"tr",{},[123,124,126],"th",{"align":125},"left","Metric",[123,128,130,133],{"align":129},"center",[24,131,29],{"href":26,"rel":132},[28]," (Rust + HIP)",[123,135,136,97,138,101],{"align":129},[48,137,50],{},[48,139,100],{},[123,141,142,97,144,101],{"align":129},[48,143,54],{},[48,145,108],{},[123,147,148],{"align":125},"Architectural Win",[150,151,152,177,201,223,243],"tbody",{},[120,153,154,160,165,168,171],{},[155,156,157],"td",{"align":125},[21,158,159],{},"Decode Throughput",[155,161,162],{"align":129},[21,163,164],{},"83.0 tok\u002Fs",[155,166,167],{"align":129},"71.3 tok\u002Fs",[155,169,170],{"align":129},"72.0 tok\u002Fs",[155,172,173,176],{"align":125},[21,174,175],{},"+16.4% vs llama.cpp"," (+15.3% vs Ollama)",[120,178,179,184,189,192,195],{},[155,180,181],{"align":125},[21,182,183],{},"Time-To-First-Token (TTFT)",[155,185,186],{"align":129},[21,187,188],{},"48.5 ms",[155,190,191],{"align":129},"121.0 ms",[155,193,194],{"align":129},"173.2 ms",[155,196,197,200],{"align":125},[21,198,199],{},"-59.9% latency reduction"," (-72.0% vs Ollama)",[120,202,203,208,213,216,218],{},[155,204,205],{"align":125},[21,206,207],{},"Prefill Kernel Launches",[155,209,210],{"align":129},[21,211,212],{},"280 launches",[155,214,215],{"align":129},"~3,300 launches",[155,217,215],{"align":129},[155,219,220],{"align":125},[21,221,222],{},"-91.5% host CPU driver dispatch queue",[120,224,225,230,235,238,240],{},[155,226,227],{"align":125},[21,228,229],{},"Streaming Frame Overhead",[155,231,232],{"align":129},[21,233,234],{},"\u003C 0.3%",[155,236,237],{"align":129},"n\u002Fa",[155,239,237],{"align":129},[155,241,242],{"align":125},"Isolated in-process A\u002FB\u002FC measurement",[120,244,245,250,255,257,259],{},[155,246,247],{"align":125},[21,248,249],{},"Weight Precision",[155,251,252],{"align":129},[21,253,254],{},"BF16",[155,256,254],{"align":129},[155,258,254],{"align":129},[155,260,261],{"align":125},"Byte-identical parameters across all 3 arms",[14,263,264,265,267,268,271,272,275],{},"Across the wider model family, ",[21,266,29],{}," maintained its decode and TTFT lead from ",[21,269,270],{},"0.8B (285.9 vs 210.6 tok\u002Fs, +35.7%)"," all the way to ",[21,273,274],{},"9B (49.1 vs 44.6 tok\u002Fs, +10.2%)",", where execution hits the 960 GB\u002Fs physical memory bandwidth ceiling of the GDDR6 bus.",[14,277,278,279,281],{},"Getting there meant debugging a 3,300-kernel launch storm that locked the CPU driver queue for 50 ms, diagnosing a phantom 600 ms stall caused by an un-flushed 8KB HTTP socket buffer, debunking a widespread myth about per-token network streaming overhead, and restructuring decode attention into a 4-way parallel split-KV reduction kernel inspired by ",[48,280,50],{}," itself.",[14,283,284],{},"The sections below cover the engineering journey, the data, the kernel code, and the caveats.",[286,287],"hr",{},[289,290,292],"h2",{"id":291},"️-the-test-bench-the-parity-rules","🛠️ The Test Bench & The Parity Rules",[14,294,295],{},"Before any speedup claims, the benchmark needs ground rules. Comparing a raw CLI binary against an HTTP daemon, or quantized weights against float16, produces meaningless marketing numbers. I enforced strict architectural parity:",[297,298,299,343,394,408,445],"ol",{},[300,301,302,305,306],"li",{},[21,303,304],{},"Hardware Configuration",":\n",[307,308,309,318,324,333],"ul",{},[300,310,311,314,315,317],{},[21,312,313],{},"GPU",": AMD Radeon RX 7900 XTX (Navi 31, RDNA3, target architecture ",[48,316,64],{},").",[300,319,320,323],{},[21,321,322],{},"Compute Units",": 96 Compute Units, 6,144 Stream Processors.",[300,325,326,329,330,89],{},[21,327,328],{},"VRAM",": 24 GB GDDR6 running on a 384-bit memory bus with ",[21,331,332],{},"960 GB\u002Fs theoretical peak bandwidth",[300,334,335,338,339,342],{},[21,336,337],{},"Software Stack",": Linux x86_64, ROCm 7.2, HIP compiler (",[48,340,341],{},"hipcc","), GCC 14.",[300,344,345,305,348],{},[21,346,347],{},"Independent Network Daemons",[307,349,350,371,381,387],{},[300,351,352,353,355,356,359,360,355,362,365,366,355,368,89],{},"Every engine ran as an independent background daemon listening on localhost: ",[21,354,29],{}," on port ",[48,357,358],{},"8003",", ",[48,361,100],{},[48,363,364],{},"8001",", and ",[48,367,108],{},[48,369,370],{},"11434",[300,372,373,374,377,378,89],{},"Test clients issued HTTP ",[48,375,376],{},"POST \u002Fv1\u002Fchat\u002Fcompletions"," requests over TCP loopback sockets with ",[48,379,380],{},"{\"stream\": true}",[300,382,383,384,386],{},"I measured ",[21,385,183],{}," from the moment the socket opened until the first SSE chunk arrived at the client.",[300,388,389,390,393],{},"I calculated ",[21,391,392],{},"decode throughput"," as total generated tokens divided by the duration of the token generation phase.",[300,395,396,305,399],{},[21,397,398],{},"Weight Parity",[307,400,401],{},[300,402,403,404,407],{},"Native unquantized ",[21,405,406],{},"bfloat16 (BF16)"," parameters across all three engines. No quantization artifacts or precision mismatches.",[300,409,410,305,413],{},[21,411,412],{},"100% ROCm GPU Acceleration Parity",[307,414,415,418,432],{},[300,416,417],{},"Neither baseline could fall back to CPU compute.",[300,419,420,421,423,424,427,428,431],{},"I compiled ",[48,422,100],{}," natively with ",[48,425,426],{},"GGML_HIP_GRAPHS=ON"," and HIPBLAS support and ran it with ",[48,429,430],{},"-ngl 999"," to offload all 32 model layers plus embeddings and heads into VRAM.",[300,433,434,436,437,440,441,444],{},[48,435,108],{}," ran on CachyOS's hardware-accelerated ",[48,438,439],{},"ollama-rocm"," package, dynamically linking ",[48,442,443],{},"\u002Fusr\u002Flib\u002Follama\u002Frocm_v7_2\u002Flibggml-hip.so"," and pinning all 34 layers (8,023.7 MiB VRAM buffer).",[300,446,447,305,450],{},[21,448,449],{},"Exactness Gate",[307,451,452],{},[300,453,454,455,458],{},"I checked greedy decode outputs token-for-token against reference completions from Hugging Face ",[48,456,457],{},"transformers"," and rejected any engine state that failed exact token-id parity. A broken kernel that skips operations runs faster because it computes the wrong thing.",[286,460],{},[289,462,464],{"id":463},"the-beast-qwen-35-hybrid-architecture","🧠 The Beast: Qwen 3.5 Hybrid Architecture",[14,466,467,468,538],{},"Standard autoregressive Transformers (such as LLaMA 3 or Mistral) use standard Multi-Head or Grouped-Query Attention across every single layer. In those architectures, the Key-Value (KV) cache grows linearly (",[469,470,473,506],"span",{"className":471},[472],"katex",[469,474,477],{"className":475},[476],"katex-mathml",[478,479,481],"math",{"xmlns":480},"http:\u002F\u002Fwww.w3.org\u002F1998\u002FMath\u002FMathML",[482,483,484,501],"semantics",{},[485,486,487,491,496,499],"mrow",{},[488,489,490],"mi",{},"O",[492,493,495],"mo",{"stretchy":494},"false","(",[488,497,498],{},"T",[492,500,101],{"stretchy":494},[502,503,505],"annotation",{"encoding":504},"application\u002Fx-tex","O(T)",[469,507,511],{"className":508,"ariaHidden":510},[509],"katex-html","true",[469,512,515,520,526,530,534],{"className":513},[514],"base",[469,516],{"className":517,"style":519},[518],"strut","height:1em;vertical-align:-0.25em;",[469,521,490],{"className":522,"style":525},[523,524],"mord","mathnormal","margin-right:0.0278em;",[469,527,495],{"className":528},[529],"mopen",[469,531,498],{"className":532,"style":533},[523,524],"margin-right:0.1389em;",[469,535,101],{"className":536},[537],"mclose",") with every generated token. By layer 32 at context length 8,192, a standard transformer allocates gigabytes of VRAM strictly to store KV activations.",[14,540,541,542,545],{},"Qwen 3.5 4B breaks this paradigm by utilizing a ",[21,543,544],{},"hybrid topology"," across its 32 layers:",[307,547,548,866],{},[300,549,550,553,554,557,558,741,742,745,746,816,817,865],{},[21,551,552],{},"24 Gated DeltaNet (GDN) Linear Attention Layers",":\nThese layers replace standard quadratic attention with a ",[21,555,556],{},"fixed-size recurrent state matrix"," ",[469,559,561,608],{"className":560},[472],[469,562,564],{"className":563},[476],[478,565,566],{"xmlns":480},[482,567,568,605],{},[485,569,570,579,582],{},[571,572,573,576],"msub",{},[488,574,575],{},"S",[488,577,578],{},"t",[492,580,581],{},"∈",[583,584,585,589],"msup",{},[488,586,588],{"mathvariant":587},"double-struck","R",[485,590,591,595,598,601,603],{},[592,593,594],"mn",{},"32",[492,596,597],{},"×",[592,599,600],{},"128",[492,602,597],{},[592,604,600],{},[502,606,607],{"encoding":504},"S_t \\in \\mathbb{R}^{32 \\times 128 \\times 128}",[469,609,611,687],{"className":610,"ariaHidden":510},[509],[469,612,614,618,675,680,684],{"className":613},[514],[469,615],{"className":616,"style":617},[518],"height:0.8333em;vertical-align:-0.15em;",[469,619,621,625],{"className":620},[523],[469,622,575],{"className":623,"style":624},[523,524],"margin-right:0.0576em;",[469,626,629],{"className":627},[628],"msupsub",[469,630,634,666],{"className":631},[632,633],"vlist-t","vlist-t2",[469,635,638,661],{"className":636},[637],"vlist-r",[469,639,643],{"className":640,"style":642},[641],"vlist","height:0.2806em;",[469,644,646,651],{"style":645},"top:-2.55em;margin-left:-0.0576em;margin-right:0.05em;",[469,647],{"className":648,"style":650},[649],"pstrut","height:2.7em;",[469,652,658],{"className":653},[654,655,656,657],"sizing","reset-size6","size3","mtight",[469,659,578],{"className":660},[523,524,657],[469,662,665],{"className":663},[664],"vlist-s","​",[469,667,669],{"className":668},[637],[469,670,673],{"className":671,"style":672},[641],"height:0.15em;",[469,674],{},[469,676],{"className":677,"style":679},[678],"mspace","margin-right:0.2778em;",[469,681,581],{"className":682},[683],"mrel",[469,685],{"className":686,"style":679},[678],[469,688,690,694],{"className":689},[514],[469,691],{"className":692,"style":693},[518],"height:0.8141em;",[469,695,697,701],{"className":696},[523],[469,698,588],{"className":699},[523,700],"mathbb",[469,702,704],{"className":703},[628],[469,705,707],{"className":706},[632],[469,708,710],{"className":709},[637],[469,711,713],{"className":712,"style":693},[641],[469,714,716,719],{"style":715},"top:-3.063em;margin-right:0.05em;",[469,717],{"className":718,"style":650},[649],[469,720,722],{"className":721},[654,655,656,657],[469,723,725,728,732,735,738],{"className":724},[523,657],[469,726,594],{"className":727},[523,657],[469,729,597],{"className":730},[731,657],"mbin",[469,733,600],{"className":734},[523,657],[469,736,597],{"className":737},[731,657],[469,739,600],{"className":740},[523,657]," in single precision (FP32). Each layer maintains a 2.0 MB recurrent state, totaling ",[21,743,744],{},"48.0 MB across all 24 layers",". As tokens arrive, ",[469,747,749,767],{"className":748},[472],[469,750,752],{"className":751},[476],[478,753,754],{"xmlns":480},[482,755,756,764],{},[485,757,758],{},[571,759,760,762],{},[488,761,575],{},[488,763,578],{},[502,765,766],{"encoding":504},"S_t",[469,768,770],{"className":769,"ariaHidden":510},[509],[469,771,773,776],{"className":772},[514],[469,774],{"className":775,"style":617},[518],[469,777,779,782],{"className":778},[523],[469,780,575],{"className":781,"style":624},[523,524],[469,783,785],{"className":784},[628],[469,786,788,808],{"className":787},[632,633],[469,789,791,805],{"className":790},[637],[469,792,794],{"className":793,"style":642},[641],[469,795,796,799],{"style":645},[469,797],{"className":798,"style":650},[649],[469,800,802],{"className":801},[654,655,656,657],[469,803,578],{"className":804},[523,524,657],[469,806,665],{"className":807},[664],[469,809,811],{"className":810},[637],[469,812,814],{"className":813,"style":672},[641],[469,815],{}," updates causally in-place via a 1D convolution and a gated delta update rule. Memory consumption is ",[21,818,819,820],{},"strictly ",[469,821,823,844],{"className":822},[472],[469,824,826],{"className":825},[476],[478,827,828],{"xmlns":480},[482,829,830,841],{},[485,831,832,834,836,839],{},[488,833,490],{},[492,835,495],{"stretchy":494},[592,837,838],{},"1",[492,840,101],{"stretchy":494},[502,842,843],{"encoding":504},"O(1)",[469,845,847],{"className":846,"ariaHidden":510},[509],[469,848,850,853,856,859,862],{"className":849},[514],[469,851],{"className":852,"style":519},[518],[469,854,490],{"className":855,"style":525},[523,524],[469,857,495],{"className":858},[529],[469,860,838],{"className":861},[523],[469,863,101],{"className":864},[537],"—it never expands by a single byte regardless of whether the prompt is 10 tokens or 10,000 tokens long.",[300,867,868,871,872,875,876,879],{},[21,869,870],{},"8 Full Grouped-Query Attention (GQA) Layers",":\nPositioned at every 4th layer (specifically layers 3, 7, 11, 15, 19, 23, 27, and 31). These 8 layers maintain a standard dynamic KV cache (",[48,873,874],{},"[4 heads, seq_len, 256 dim]"," in BF16), consuming ",[21,877,878],{},"32 KB per token across all 8 layers",". These full attention layers preserve exact context retrieval and multi-needle recall over long sequences.",[881,882,883],"blockquote",{},[14,884,885,888,891],{},[469,886,887],{},"!NOTE",[21,889,890],{},"The Detective's Notebook Analogy",":\nImagine a detective investigating a complex case. For 75% of their daily work (the 24 GDN layers), the detective writes condensed, rolling summaries into a small, pocket-sized notebook of fixed size. The notebook never gets heavier, and older facts are continuously merged and compressed. For the remaining 25% of critical evidence (the 8 Full Attention layers), the detective keeps verbatim transcripts of witness testimonies, meticulously re-reads every transcript from page one whenever a new question arrives.",[14,893,894],{},"This hybrid structure dictates the inference engine's performance profile: memory allocation is lean, but the execution pipeline constantly alternates between linear recurrent scans and quadratic attention projections.",[14,896,897],{},[40,898],{"alt":899,"src":900,"height":901,"width":902},"Qwen 3.5 layer map: 24 Gated DeltaNet layers with fixed 48 MB state and 8 full-attention layers with a growing KV cache | wide","\u002Fimages\u002Fblog\u002Fstrata-hybrid-layers.svg#wide",600,1200,[286,904],{},[289,906,908],{"id":907},"act-0-from-317-to-822-toks-the-raw-decode-engine","Act 0: From 31.7 to 82.2 tok\u002Fs (The Raw Decode Engine)",[14,910,911,912,915,916,918],{},"My first naive port of the Qwen 3.5 architecture in Rust and raw HIP ran at ",[21,913,914],{},"31.7 tok\u002Fs",". It produced correct output but trailed ",[48,917,50],{}," (71.3 tok\u002Fs) by more than 55%.",[14,920,921],{},"Reaching 82+ tok\u002Fs took four optimization passes on the GPU kernels:",[923,924,929],"pre",{"className":925,"code":927,"language":928},[926],"language-text","Pass 0: Baseline naive port                       31.7 tok\u002Fs\nPass 1: Buffer reuse & sync elimination          44.3 tok\u002Fs (+40%)\nPass 2: Projection GEMM fusing                    49.2 tok\u002Fs (+11%)\nPass 3: HIP Graphs + Vectorized GEMV + Argmax    81.2 tok\u002Fs (+65%)\nPass 4: GDN Thread Block Occupancy Tuning         82.2 tok\u002Fs (+1.2%)\n","text",[48,930,927],{"__ignoreMap":931},"",[14,933,934],{},[40,935],{"alt":936,"src":937,"height":938,"width":902},"Decode throughput climbing from 31.7 to 82.2 tok\u002Fs across four optimization passes, against the llama.cpp baseline of 71.3 | wide","\u002Fimages\u002Fblog\u002Fstrata-decode-waterfall.svg#wide",650,[940,941,943],"h3",{"id":942},"why-safe-rust-with-hip","Why Safe Rust with HIP?",[14,945,946,947,949],{},"Writing bare-metal GPU kernels for AMD RDNA3 requires compiling HIP C++ through ",[48,948,341],{},". The host-side architecture also matters for low-latency local AI serving. I wrote Strata's host engine in Rust, which gave three benefits:",[297,951,952,965,971],{},[300,953,954,957,958,317],{},[21,955,956],{},"FFI Encapsulation",": I isolated unsafe raw device pointers and HIP runtime invocations in a minimal, audited FFI module (",[24,959,962],{"href":960,"rel":961},"https:\u002F\u002Fgithub.com\u002Fmihailtd\u002Fstrata\u002Fblob\u002Fmain\u002Fapps\u002Fruntime-next\u002Fsrc\u002Fhip.rs",[28],[48,963,964],{},"hip.rs",[300,966,967,970],{},[21,968,969],{},"Safe Graph Orchestration",": I wrote the entire model execution DAG, KV management, and network server in 100% safe Rust.",[300,972,973,976],{},[21,974,975],{},"Zero Interpreter Latency",": Leaving Python removed the Global Interpreter Lock (GIL), garbage collection pauses, and runtime dispatch overhead, and these hurt real-time local LLM applications.",[940,978,980],{"id":979},"the-optimization-passes","The Optimization Passes",[307,982,983,993,1009,1164],{},[300,984,985,988,989,992],{},[21,986,987],{},"Pass 1 (31.7 → 44.3 tok\u002Fs)",": We eliminated redundant host-to-device synchronizations (",[48,990,991],{},"hipDeviceSynchronize",") between consecutive layers and replaced transient VRAM allocations with pre-allocated static execution scratchpads.",[300,994,995,998,999,1002,1003,1006,1007,89],{},[21,996,997],{},"Pass 2 (44.3 → 49.2 tok\u002Fs)",": We fused projection operations, reading directly from combined GEMM output buffers. We evaluated ",[48,1000,1001],{},"hipBLASLt"," as an alternative to ",[48,1004,1005],{},"hipblasGemmEx",", but benchmarks showed it ran marginally slower on Qwen's specific rectangular matrix dimensions, so we retained ",[48,1008,1005],{},[300,1010,1011,1014,1015],{},[21,1012,1013],{},"Pass 3 (49.2 → 81.2 tok\u002Fs)",": This was the architectural breakthrough, unlocked by three changes:\n",[297,1016,1017,1084,1147],{},[300,1018,1019,1022,1023,1079,1080,1083],{},[21,1020,1021],{},"HIP Graph Replay",": Autoregressive decode evaluates exactly one token (",[469,1024,1026,1046],{"className":1025},[472],[469,1027,1029],{"className":1028},[476],[478,1030,1031],{"xmlns":480},[482,1032,1033,1043],{},[485,1034,1035,1038,1041],{},[488,1036,1037],{},"M",[492,1039,1040],{},"=",[592,1042,838],{},[502,1044,1045],{"encoding":504},"M=1",[469,1047,1049,1069],{"className":1048,"ariaHidden":510},[509],[469,1050,1052,1056,1060,1063,1066],{"className":1051},[514],[469,1053],{"className":1054,"style":1055},[518],"height:0.6833em;",[469,1057,1037],{"className":1058,"style":1059},[523,524],"margin-right:0.109em;",[469,1061],{"className":1062,"style":679},[678],[469,1064,1040],{"className":1065},[683],[469,1067],{"className":1068,"style":679},[678],[469,1070,1072,1076],{"className":1071},[514],[469,1073],{"className":1074,"style":1075},[518],"height:0.6444em;",[469,1077,838],{"className":1078},[523],") per step. Re-issuing dozens of small kernels every 12 milliseconds flooded the host CPU with dispatch work. By capturing the complete decode iteration into a frozen ",[48,1081,1082],{},"hipGraphExec_t",", host dispatch cost dropped to virtually zero.",[300,1085,1086,1092,1093,1143,1144,1146],{},[21,1087,1088,1089,101],{},"Custom Vectorized GEMV (",[48,1090,1091],{},"ushort4",": Single-token decode is a matrix-vector product, not a general matrix-matrix multiply. Standard GEMM kernels use warp tiles poorly at ",[469,1094,1096,1113],{"className":1095},[472],[469,1097,1099],{"className":1098},[476],[478,1100,1101],{"xmlns":480},[482,1102,1103,1111],{},[485,1104,1105,1107,1109],{},[488,1106,1037],{},[492,1108,1040],{},[592,1110,838],{},[502,1112,1045],{"encoding":504},[469,1114,1116,1134],{"className":1115,"ariaHidden":510},[509],[469,1117,1119,1122,1125,1128,1131],{"className":1118},[514],[469,1120],{"className":1121,"style":1055},[518],[469,1123,1037],{"className":1124,"style":1059},[523,524],[469,1126],{"className":1127,"style":679},[678],[469,1129,1040],{"className":1130},[683],[469,1132],{"className":1133,"style":679},[678],[469,1135,1137,1140],{"className":1136},[514],[469,1138],{"className":1139,"style":1075},[518],[469,1141,838],{"className":1142},[523],". We implemented a custom HIP GEMV kernel using vectorized 64-bit loads (",[48,1145,1091],{},", loading 4 BF16 elements per instruction), doubling global memory throughput on the dominant projection matrices.",[300,1148,1149,1152,1153,1160,1161,89],{},[21,1150,1151],{},"On-Device Argmax Reduction",": In naive implementations, the host copies the final logit tensor (248,320 floating-point numbers) over PCIe to the CPU, which performs an argmax to select the next token. We replaced this with an on-device parallel reduction kernel (",[24,1154,1157],{"href":1155,"rel":1156},"https:\u002F\u002Fgithub.com\u002Fmihailtd\u002Fstrata\u002Fblob\u002Fmain\u002Fapps\u002Fruntime-next\u002Fsrc\u002Fkernels\u002Fargmax.hip",[28],[48,1158,1159],{},"argmax.hip","), reducing host-bound PCIe traffic from ",[21,1162,1163],{},"993 KB per token to a single 4-byte integer",[300,1165,1166,1169,1170,1173,1174,1177,1178,89],{},[21,1167,1168],{},"Pass 4 (81.2 → 82.2 tok\u002Fs)",": We profiled kernel execution using AMD's official profiler, ",[48,1171,1172],{},"rocprofv3",". The trace revealed that ",[48,1175,1176],{},"gdn_recurrent_decode"," used only 32 of the RX 7900 XTX's 96 Compute Units (launching one workgroup per attention head across 32 heads). Increasing the thread block size from 128 to 1,024 threads improved wave occupancy on Navi 31, reducing kernel execution from ",[21,1179,1180],{},"60.5 μs to 24.9 μs",[14,1182,1183],{},"With decode reaching 82.2 tok\u002Fs in internal harnesses, we turned to real-world prompt prefill and HTTP serving. That is where things broke.",[286,1185],{},[289,1187,1189],{"id":1188},"act-1-the-3300-kernel-launch-storm","Act 1: The 3,300-Kernel Launch Storm",[14,1191,1192,1193,1221],{},"When we first ported prompt prefill, the forward pass processed incoming tokens sequentially inside a loop over sequence length ",[469,1194,1196,1209],{"className":1195},[472],[469,1197,1199],{"className":1198},[476],[478,1200,1201],{"xmlns":480},[482,1202,1203,1207],{},[485,1204,1205],{},[488,1206,498],{},[502,1208,498],{"encoding":504},[469,1210,1212],{"className":1211,"ariaHidden":510},[509],[469,1213,1215,1218],{"className":1214},[514],[469,1216],{"className":1217,"style":1055},[518],[469,1219,498],{"className":1220,"style":533},[523,524],":",[307,1223,1224,1264,1299,1336],{},[300,1225,1226,1229,1230,1260,1261],{},[21,1227,1228],{},"Causal Conv1D",": 54 tokens ",[469,1231,1233,1247],{"className":1232},[472],[469,1234,1236],{"className":1235},[476],[478,1237,1238],{"xmlns":480},[482,1239,1240,1244],{},[485,1241,1242],{},[492,1243,597],{},[502,1245,1246],{"encoding":504},"\\times",[469,1248,1250],{"className":1249,"ariaHidden":510},[509],[469,1251,1253,1257],{"className":1252},[514],[469,1254],{"className":1255,"style":1256},[518],"height:0.6667em;vertical-align:-0.0833em;",[469,1258,597],{"className":1259},[523]," 24 GDN layers = ",[21,1262,1263],{},"1,296 kernel launches",[300,1265,1266,1229,1269,1260,1297],{},[21,1267,1268],{},"GDN Gate & Beta Computation",[469,1270,1272,1285],{"className":1271},[472],[469,1273,1275],{"className":1274},[476],[478,1276,1277],{"xmlns":480},[482,1278,1279,1283],{},[485,1280,1281],{},[492,1282,597],{},[502,1284,1246],{"encoding":504},[469,1286,1288],{"className":1287,"ariaHidden":510},[509],[469,1289,1291,1294],{"className":1290},[514],[469,1292],{"className":1293,"style":1256},[518],[469,1295,597],{"className":1296},[523],[21,1298,1263],{},[300,1300,1301,1229,1304,1332,1333],{},[21,1302,1303],{},"RoPE Position Embeddings",[469,1305,1307,1320],{"className":1306},[472],[469,1308,1310],{"className":1309},[476],[478,1311,1312],{"xmlns":480},[482,1313,1314,1318],{},[485,1315,1316],{},[492,1317,597],{},[502,1319,1246],{"encoding":504},[469,1321,1323],{"className":1322,"ariaHidden":510},[509],[469,1324,1326,1329],{"className":1325},[514],[469,1327],{"className":1328,"style":1256},[518],[469,1330,597],{"className":1331},[523]," 8 Attention layers = ",[21,1334,1335],{},"432 kernel launches",[300,1337,1338,1229,1341,1332,1369],{},[21,1339,1340],{},"KV Cache Appends",[469,1342,1344,1357],{"className":1343},[472],[469,1345,1347],{"className":1346},[476],[478,1348,1349],{"xmlns":480},[482,1350,1351,1355],{},[485,1352,1353],{},[492,1354,597],{},[502,1356,1246],{"encoding":504},[469,1358,1360],{"className":1359,"ariaHidden":510},[509],[469,1361,1363,1366],{"className":1362},[514],[469,1364],{"className":1365,"style":1256},[518],[469,1367,597],{"className":1368},[523],[21,1370,1335],{},[14,1372,1373,1374,1377,1378,89],{},"A standard ",[21,1375,1376],{},"54-token prompt"," triggered ",[21,1379,1380],{},"~3,300 distinct HIP kernel launches",[881,1382,1383],{},[14,1384,1385,1387,1390],{},[469,1386,887],{},[21,1388,1389],{},"The Chef and the Grains of Rice",":\nImagine a state-of-the-art commercial kitchen staffed by 6,144 eager cooks (the GPU stream processors). Instead of bringing in a sack of rice to cook, an assistant walks through the kitchen door 3,300 times in a row, handing the cooks exactly one grain of rice per trip. The cooks spend 95% of their time staring at the swinging kitchen door waiting for the assistant to walk through.",[14,1392,1393,1394,1397],{},"At 15–20 μs of driver dispatch latency per launch on Linux ROCm, host CPU driver overhead consumed ",[21,1395,1396],{},"over 50 milliseconds"," before the GPU finished computing the prompt.",[940,1399,1401],{"id":1400},"the-fix-batched-chunk-kernels","The Fix: Batched Chunk Kernels",[14,1403,1404],{},"We eliminated the per-token dispatch loops by writing four batched prefill kernels that process the entire prompt sequence in a single launch per layer:",[114,1406,1407,1423],{},[117,1408,1409],{},[120,1410,1411,1414,1417,1420],{},[123,1412,1413],{"align":125},"Kernel Operation",[123,1415,1416],{"align":129},"Unbatched Launches (T=54)",[123,1418,1419],{"align":129},"Batched Launches",[123,1421,1422],{"align":125},"Parallelization Mechanism",[150,1424,1425,1527,1611,1629,1645],{},[120,1426,1427,1431,1434,1439],{},[155,1428,1429],{"align":125},[21,1430,1228],{},[155,1432,1433],{"align":129},"1,296",[155,1435,1436],{"align":129},[21,1437,1438],{},"24",[155,1440,1441,1442,1497,1498,1526],{"align":125},"Bounded lookback (",[469,1443,1445,1465],{"className":1444},[472],[469,1446,1448],{"className":1447},[476],[478,1449,1450],{"xmlns":480},[482,1451,1452,1462],{},[485,1453,1454,1457,1459],{},[488,1455,1456],{},"k",[492,1458,1040],{},[592,1460,1461],{},"4",[502,1463,1464],{"encoding":504},"k=4",[469,1466,1468,1488],{"className":1467,"ariaHidden":510},[509],[469,1469,1471,1475,1479,1482,1485],{"className":1470},[514],[469,1472],{"className":1473,"style":1474},[518],"height:0.6944em;",[469,1476,1456],{"className":1477,"style":1478},[523,524],"margin-right:0.0315em;",[469,1480],{"className":1481,"style":679},[678],[469,1483,1040],{"className":1484},[683],[469,1486],{"className":1487,"style":679},[678],[469,1489,1491,1494],{"className":1490},[514],[469,1492],{"className":1493,"style":1075},[518],[469,1495,1461],{"className":1496},[523],"): 1 thread per channel scans all ",[469,1499,1501,1514],{"className":1500},[472],[469,1502,1504],{"className":1503},[476],[478,1505,1506],{"xmlns":480},[482,1507,1508,1512],{},[485,1509,1510],{},[488,1511,498],{},[502,1513,498],{"encoding":504},[469,1515,1517],{"className":1516,"ariaHidden":510},[509],[469,1518,1520,1523],{"className":1519},[514],[469,1521],{"className":1522,"style":1055},[518],[469,1524,498],{"className":1525,"style":533},[523,524]," tokens carrying a 4-element register window.",[120,1528,1529,1534,1536,1540],{},[155,1530,1531],{"align":125},[21,1532,1533],{},"GDN Gate & Beta",[155,1535,1433],{"align":129},[155,1537,1538],{"align":129},[21,1539,1438],{},[155,1541,1542,1543,1607,1608,89],{"align":125},"2D parallel grid over ",[469,1544,1546,1572],{"className":1545},[472],[469,1547,1549],{"className":1548},[476],[478,1550,1551],{"xmlns":480},[482,1552,1553,1569],{},[485,1554,1555,1557,1561,1564,1567],{},[492,1556,495],{"stretchy":494},[1558,1559,1560],"mtext",{},"token",[492,1562,1563],{"separator":510},",",[1558,1565,1566],{},"head",[492,1568,101],{"stretchy":494},[502,1570,1571],{"encoding":504},"(\\text{token}, \\text{head})",[469,1573,1575],{"className":1574,"ariaHidden":510},[509],[469,1576,1578,1581,1584,1590,1594,1598,1604],{"className":1577},[514],[469,1579],{"className":1580,"style":519},[518],[469,1582,495],{"className":1583},[529],[469,1585,1587],{"className":1586},[523,928],[469,1588,1560],{"className":1589},[523],[469,1591,1563],{"className":1592},[1593],"mpunct",[469,1595],{"className":1596,"style":1597},[678],"margin-right:0.1667em;",[469,1599,1601],{"className":1600},[523,928],[469,1602,1566],{"className":1603},[523],[469,1605,101],{"className":1606},[537],": Zero cross-token recurrence; heads indexed via ",[48,1609,1610],{},"tid % num_v_heads",[120,1612,1613,1618,1621,1626],{},[155,1614,1615],{"align":125},[21,1616,1617],{},"RoPE Embeddings",[155,1619,1620],{"align":129},"432",[155,1622,1623],{"align":129},[21,1624,1625],{},"8",[155,1627,1628],{"align":125},"Vectorized position indexing: Each token rotates via its precomputed absolute position buffer.",[120,1630,1631,1636,1638,1642],{},[155,1632,1633],{"align":125},[21,1634,1635],{},"KV Cache Append",[155,1637,1620],{"align":129},[155,1639,1640],{"align":129},[21,1641,1625],{},[155,1643,1644],{"align":125},"Direct strided scatter: Each token writes directly into its designated memory offset in VRAM.",[120,1646,1647,1652,1657,1662],{},[155,1648,1649],{"align":125},[21,1650,1651],{},"Total Host Launches",[155,1653,1654],{"align":129},[21,1655,1656],{},"~3,300",[155,1658,1659],{"align":129},[21,1660,1661],{},"280",[155,1663,1664],{"align":125},[21,1665,1666],{},"-91.5% reduction in CPU driver dispatch queue lag",[14,1668,1669],{},[40,1670],{"alt":1671,"src":1672,"height":1673,"width":902},"Kernel launches per prefill: per-token loops versus batched kernels, 3,300 down to 280 | wide","\u002Fimages\u002Fblog\u002Fstrata-launch-storm.svg#wide",700,[14,1675,1676,1677,1680],{},"Batching these operations reduced internal GPU prefill latency from ",[21,1678,1679],{},"78.95 ms to 71.28 ms",". We expected HTTP Time-To-First-Token to drop proportionally.",[14,1682,1683],{},"Instead, TTFT stalled.",[286,1685],{},[289,1687,1689],{"id":1688},"act-2-the-phantom-stall-a-600ms-bug-behind-a-70ms-kernel","Act 2: The Phantom Stall: A 600ms Bug Behind a 70ms Kernel",[14,1691,1692,1693,1696,1697,89],{},"When we pointed ",[48,1694,1695],{},"curl"," and Python clients at our HTTP endpoint, Time-To-First-Token sat at ",[21,1698,1699],{},"613–640 milliseconds",[14,1701,1702],{},"Our GPU forward pass was taking 71 milliseconds. Where were the other 540 milliseconds going?",[14,1704,1705],{},"We systematically audited every component of the serving pipeline:",[297,1707,1708,1718,1731],{},[300,1709,1710,1713,1714,1717],{},[21,1711,1712],{},"Client Parsing Overhead",": We tested with raw ",[48,1715,1716],{},"curl -N -w \"%{time_starttransfer}\\n\" -s -o \u002Fdev\u002Fnull",", confirming the same 600+ ms stall on bare sockets.",[300,1719,1720,1723,1724,1727,1728,89],{},[21,1721,1722],{},"State Reset",": We profiled ",[48,1725,1726],{},"engine.reset_state()",", which cleared recurrent buffers in ",[21,1729,1730],{},"0.05 ms",[300,1732,1733,1736,1737,1740],{},[21,1734,1735],{},"BPE Tokenization",": We benchmarked the Hugging Face tokenizer decoding 350 tokens: ",[21,1738,1739],{},"3.32 ms total"," (~9 μs per token).",[940,1742,1744],{"id":1743},"the-smoking-gun-the-un-flushed-8kb-buffer","The Smoking Gun: The Un-Flushed 8KB Buffer",[14,1746,1747,1748,1751],{},"We dug into our HTTP server dependency: ",[48,1749,1750],{},"tiny_http"," (version 0.12.0).",[14,1753,1754,1755,1757,1758,1761,1762,1765,1766,1769],{},"In ",[48,1756,1750],{},", the standard streaming API uses ",[48,1759,1760],{},"Response::raw_print",", which internally wraps the HTTP response body in ",[48,1763,1764],{},"chunked_transfer::Encoder",". Inspecting the source of ",[48,1767,1768],{},"chunked_transfer"," revealed a default that explained the stall:",[923,1771,1775],{"className":1772,"code":1773,"language":1774,"meta":931,"style":931},"language-rust shiki shiki-themes github-light github-dark","\u002F\u002F Inside chunked_transfer::Encoder\npub struct Encoder\u003CW> {\n    writer: W,\n    buffer: [u8; 8192], \u002F\u002F \u003C-- HARDCODED 8KB INTERNAL BUFFER\n    buffer_len: usize,\n}\n","rust",[48,1776,1777,1784,1790,1796,1802,1808],{"__ignoreMap":931},[469,1778,1781],{"class":1779,"line":1780},"line",1,[469,1782,1783],{},"\u002F\u002F Inside chunked_transfer::Encoder\n",[469,1785,1787],{"class":1779,"line":1786},2,[469,1788,1789],{},"pub struct Encoder\u003CW> {\n",[469,1791,1793],{"class":1779,"line":1792},3,[469,1794,1795],{},"    writer: W,\n",[469,1797,1799],{"class":1779,"line":1798},4,[469,1800,1801],{},"    buffer: [u8; 8192], \u002F\u002F \u003C-- HARDCODED 8KB INTERNAL BUFFER\n",[469,1803,1805],{"class":1779,"line":1804},5,[469,1806,1807],{},"    buffer_len: usize,\n",[469,1809,1811],{"class":1779,"line":1810},6,[469,1812,1813],{},"}\n",[14,1815,1816,1817,1820,1821,89],{},"The encoder accumulated data into an ",[21,1818,1819],{},"8,192-byte internal buffer"," and did ",[21,1822,1823],{},"not flush until the stream closed",[14,1825,1826,1827,1830,1831,1834],{},"At ~150 bytes per Server-Sent Event (SSE) JSON chunk (",[48,1828,1829],{},"data: {\"choices\":[{\"delta\":{\"content\":\"foo\"}}]}\\n\\n","), the server held ",[21,1832,1833],{},"55 generated tokens in memory"," before sending a single TCP packet.",[881,1836,1837],{},[14,1838,1839,1841,1844,1845,1848],{},[469,1840,887],{},[21,1842,1843],{},"The Reluctant Mail Carrier",":\nA mail carrier establishes an arbitrary rule: ",[17,1846,1847],{},"\"I refuse to walk down the driveway until my mailbag weighs at least 8 kilograms.\""," Even though you wrote and sealed your first letter in 70 milliseconds, the recipient has to wait while you write 54 more letters just to fill the carrier's bag.",[14,1850,1851],{},"The client was waiting for token 55, not the first token.",[14,1853,1854],{},[40,1855],{"alt":1856,"src":1857,"height":1858,"width":902},"Timeline of the 613 ms time-to-first-token stall caused by the 8 KB buffer, and the 48.5 ms result after the fix | wide","\u002Fimages\u002Fblog\u002Fstrata-ttft-stall.svg#wide",710,[940,1860,1862],{"id":1861},"the-resolution","The Resolution",[14,1864,1865,1866,1869,1870,1872,1873,1876],{},"We bypassed the buffered ",[48,1867,1868],{},"Response"," abstraction entirely with ",[48,1871,1750],{},"'s escape hatch, ",[48,1874,1875],{},"Request::into_writer()",", which gives raw access to the underlying TCP socket stream.",[14,1878,1879,1880,1883],{},"We implemented custom HTTP\u002F1.1 chunked framing with an explicit, unconditional ",[48,1881,1882],{},".flush()"," executed immediately after every SSE token:",[923,1885,1887],{"className":1772,"code":1886,"language":1774,"meta":931,"style":931},"\u002F\u002F Direct socket writer bypassing chunked_transfer 8KB buffer\nlet mut writer = request.into_writer();\nwrite!(writer, \"HTTP\u002F1.1 200 OK\\r\\nTransfer-Encoding: chunked\\r\\nContent-Type: text\u002Fevent-stream\\r\\n\\r\\n\")?;\nwriter.flush()?;\n\nfor token_str in engine.stream_tokens() {\n    let sse_payload = format!(\"data: {}\\n\\n\", serde_json::to_string(&chunk)?);\n    write!(writer, \"{:X}\\r\\n{}\\r\\n\", sse_payload.len(), sse_payload)?;\n    writer.flush()?; \u002F\u002F \u003C-- FORCES IMMEDIATE TCP PACKET EMISSION\n}\n",[48,1888,1889,1894,1899,1904,1909,1915,1920,1926,1932,1938],{"__ignoreMap":931},[469,1890,1891],{"class":1779,"line":1780},[469,1892,1893],{},"\u002F\u002F Direct socket writer bypassing chunked_transfer 8KB buffer\n",[469,1895,1896],{"class":1779,"line":1786},[469,1897,1898],{},"let mut writer = request.into_writer();\n",[469,1900,1901],{"class":1779,"line":1792},[469,1902,1903],{},"write!(writer, \"HTTP\u002F1.1 200 OK\\r\\nTransfer-Encoding: chunked\\r\\nContent-Type: text\u002Fevent-stream\\r\\n\\r\\n\")?;\n",[469,1905,1906],{"class":1779,"line":1798},[469,1907,1908],{},"writer.flush()?;\n",[469,1910,1911],{"class":1779,"line":1804},[469,1912,1914],{"emptyLinePlaceholder":1913},true,"\n",[469,1916,1917],{"class":1779,"line":1810},[469,1918,1919],{},"for token_str in engine.stream_tokens() {\n",[469,1921,1923],{"class":1779,"line":1922},7,[469,1924,1925],{},"    let sse_payload = format!(\"data: {}\\n\\n\", serde_json::to_string(&chunk)?);\n",[469,1927,1929],{"class":1779,"line":1928},8,[469,1930,1931],{},"    write!(writer, \"{:X}\\r\\n{}\\r\\n\", sse_payload.len(), sse_payload)?;\n",[469,1933,1935],{"class":1779,"line":1934},9,[469,1936,1937],{},"    writer.flush()?; \u002F\u002F \u003C-- FORCES IMMEDIATE TCP PACKET EMISSION\n",[469,1939,1941],{"class":1779,"line":1940},10,[469,1942,1813],{},[14,1944,1945],{},"The effect showed immediately:",[114,1947,1948,1964],{},[117,1949,1950],{},[120,1951,1952,1955,1958,1961],{},[123,1953,1954],{"align":125},"Engine State",[123,1956,1957],{"align":129},"TTFT (Cold Start)",[123,1959,1960],{"align":129},"TTFT (Warm Cache)",[123,1962,1963],{"align":125},"Delivery Mode",[150,1965,1966,1982],{},[120,1967,1968,1973,1976,1979],{},[155,1969,1970],{"align":125},[21,1971,1972],{},"Before Fix",[155,1974,1975],{"align":129},"700.0 ms",[155,1977,1978],{"align":129},"410.0 ms",[155,1980,1981],{"align":125},"Buffered (8KB Socket Stall)",[120,1983,1984,1989,1994,1998],{},[155,1985,1986],{"align":125},[21,1987,1988],{},"After Fix",[155,1990,1991],{"align":129},[21,1992,1993],{},"180.0 ms",[155,1995,1996],{"align":129},[21,1997,188],{},[155,1999,2000],{"align":125},[21,2001,2002],{},"Direct Unbuffered SSE Flush",[14,2004,2005,2006,2008],{},"Bypassing the socket buffer dropped warm TTFT from 410 ms to ",[21,2007,188],{},"—an 8.4x improvement that exposed the GPU's real speed.",[286,2010],{},[289,2012,2014],{"id":2013},"act-3-debunking-the-streaming-overhead-myth","Act 3: Debunking the Streaming Overhead Myth",[14,2016,2017,2018,89],{},"With the socket stall resolved, 350-token streaming benchmarks reported ",[21,2019,2020],{},"78.7–80.2 tok\u002Fs",[14,2022,2023,2024],{},"Earlier microbenchmarks on short 5-token prompts had clocked 90+ tok\u002Fs, which raised a question: ",[17,2025,2026],{},"Does flushing every single token over TCP add socket overhead that starves the GPU?",[14,2028,2029],{},"We tested two common network adjustments:",[297,2031,2032,2048],{},[300,2033,2034,2040,2041,2044,2045,89],{},[21,2035,2036,2037,101],{},"Flush Coalescing (",[48,2038,2039],{},"FLUSH_INTERVAL",": Buffering token writes into a 250 ms time window raised decode throughput to ",[21,2042,2043],{},"85.4 tok\u002Fs",", but inflated TTFT back to ",[21,2046,2047],{},"336.4 ms",[300,2049,2050,2055,2056,2058],{},[21,2051,2052],{},[48,2053,2054],{},"TCP_NODELAY",": Enabling ",[48,2057,2054],{}," on the server socket produced no measurable change (79.0 tok\u002Fs).",[940,2060,2062],{"id":2061},"the-in-process-abc-isolation-test","The In-Process A\u002FB\u002FC Isolation Test",[14,2064,2065],{},"Instead of accepting flush coalescing's trade-off, we built an in-process diagnostic that isolates the microsecond cost of each stage in the generation loop:",[307,2067,2068,2078,2084],{},[300,2069,2070,2073,2074,2077],{},[21,2071,2072],{},"Arm A",": Raw ",[48,2075,2076],{},"engine.step()"," alone (pure GPU forward pass + device argmax).",[300,2079,2080,2083],{},[21,2081,2082],{},"Arm B",": Arm A + BPE tokenizer decode + JSON string formatting + SSE chunk serialization.",[300,2085,2086,2089,2090,89],{},[21,2087,2088],{},"Arm C",": Arm B + writing to socket buffer + explicit ",[48,2091,1882],{},[14,2093,2094],{},"We executed all three arms across identical prompt workloads:",[923,2096,2099],{"className":2097,"code":2098,"language":928},[926],"Arm A (Pure GPU step):                     79.40 tok\u002Fs\nArm B (GPU + Tokenizer + JSON\u002FSSE):         80.15 tok\u002Fs\nArm C (GPU + Tokenizer + SSE + TCP Flush):  79.51 tok\u002Fs\n",[48,2100,2098],{"__ignoreMap":931},[14,2102,2103,2104,89],{},"All three arms were identical to within ",[21,2105,2106],{},"0.3%",[881,2108,2109,2117],{},[14,2110,2111,2113,2116],{},[469,2112,887],{},[21,2114,2115],{},"The Car and the Turn Signal",":\nBlaming per-token HTTP flushing for GPU decode slowdown was like blaming the flashing turn signal on your dashboard for your car losing engine horsepower. The two systems are mechanically decoupled.",[14,2118,2119,2120,2123,2124,2127],{},"Writing a 150-byte JSON chunk and flushing a TCP socket takes under ",[21,2121,2122],{},"12 microseconds"," on a modern CPU core. A GPU decode step takes ",[21,2125,2126],{},"12,000 microseconds",". The GPU was not waiting on the network.",[14,2129,2130],{},"The isolation test cleared per-token flushing. We discarded flush coalescing and kept immediate streaming.",[14,2132,2133],{},"Why, then, was throughput dropping from 80.5 tok\u002Fs down to 77.6 tok\u002Fs over long sequences?",[286,2135],{},[289,2137,2139],{"id":2138},"act-4-the-kv-cache-bottleneck-4-way-split-kv-attention","Act 4: The KV Cache Bottleneck & 4-Way Split-KV Attention",[14,2141,2142],{},"The clue lay in Arm A's segment logs. When we examined decode speed in 50-token windows across generation depth, a clear pattern emerged:",[923,2144,2147],{"className":2145,"code":2146,"language":928},[926],"Tokens  54 → 104:  80.00 tok\u002Fs\nTokens 104 → 154:  80.57 tok\u002Fs\nTokens 154 → 204:  80.38 tok\u002Fs\nTokens 204 → 254:  79.53 tok\u002Fs\nTokens 254 → 304:  77.61 tok\u002Fs  \u003C-- THROUGHPUT DECLINE\nTokens 304 → 354:  78.43 tok\u002Fs\n",[48,2148,2146],{"__ignoreMap":931},[14,2150,2151],{},"Throughput decayed steadily as sequence length grew.",[14,2153,2154],{},"I had taken the earlier 90+ tok\u002Fs measurements on 5-token prompts at positions 5–28. Over realistic 350-token generation trajectories, attention cost grew with sequence depth.",[940,2156,2158],{"id":2157},"the-mechanism","The Mechanism",[14,2160,2161],{},"In Qwen 3.5's 8 full-attention layers, each generated token must compute dot-product attention against all preceding tokens stored in the KV cache.",[14,2163,2164,2165,2172,2173,97,2176,2179,2180,97,2214,2292],{},"Our initial decode attention kernel (",[24,2166,2169],{"href":2167,"rel":2168},"https:\u002F\u002Fgithub.com\u002Fmihailtd\u002Fstrata\u002Fblob\u002Fmain\u002Fapps\u002Fruntime-next\u002Fsrc\u002Fkernels\u002Fattention.hip",[28],[48,2170,2171],{},"attention.hip",") assigned ",[21,2174,2175],{},"one thread per output dimension",[48,2177,2178],{},"head_dim = 256","). In the final weighted-V-sum reduction phase, that single thread executed a ",[21,2181,2182,2183,2213],{},"fully serial loop over all ",[469,2184,2186,2200],{"className":2185},[472],[469,2187,2189],{"className":2188},[476],[478,2190,2191],{"xmlns":480},[482,2192,2193,2198],{},[485,2194,2195],{},[488,2196,2197],{},"K",[502,2199,2197],{"encoding":504},[469,2201,2203],{"className":2202,"ariaHidden":510},[509],[469,2204,2206,2209],{"className":2205},[514],[469,2207],{"className":2208,"style":1055},[518],[469,2210,2197],{"className":2211,"style":2212},[523,524],"margin-right:0.0715em;"," past positions in the cache",[469,2215,2217,2253],{"className":2216},[472],[469,2218,2220],{"className":2219},[476],[478,2221,2222],{"xmlns":480},[482,2223,2224,2250],{},[485,2225,2226,2228,2230,2232,2235,2239,2242,2245,2248],{},[488,2227,490],{},[492,2229,495],{"stretchy":494},[488,2231,1456],{},[488,2233,2234],{},"v",[488,2236,2238],{"mathvariant":2237},"normal","_",[488,2240,2241],{},"l",[488,2243,2244],{},"e",[488,2246,2247],{},"n",[492,2249,101],{"stretchy":494},[502,2251,2252],{"encoding":504},"O(kv\\_len)",[469,2254,2256],{"className":2255,"ariaHidden":510},[509],[469,2257,2259,2263,2266,2269,2272,2276,2279,2283,2286,2289],{"className":2258},[514],[469,2260],{"className":2261,"style":2262},[518],"height:1.06em;vertical-align:-0.31em;",[469,2264,490],{"className":2265,"style":525},[523,524],[469,2267,495],{"className":2268},[529],[469,2270,1456],{"className":2271,"style":1478},[523,524],[469,2273,2234],{"className":2274,"style":2275},[523,524],"margin-right:0.0359em;",[469,2277,2238],{"className":2278,"style":525},[523],[469,2280,2241],{"className":2281,"style":2282},[523,524],"margin-right:0.0197em;",[469,2284,2244],{"className":2285},[523,524],[469,2287,2247],{"className":2288},[523,524],[469,2290,101],{"className":2291},[537]," serial scan).",[881,2294,2295],{},[14,2296,2297,2299,2302],{},[469,2298,887],{},[21,2300,2301],{},"The Lone Archivist",":\nImagine a researcher tasked with reviewing case files. When the archive holds 20 files, one person finishes quickly. When the archive expands to 350 files, that same lone researcher must read through all 350 files sequentially from start to finish, while 1,000 colleagues sit idle at their desks.",[940,2304,2306],{"id":2305},"_4-way-split-kv-parallel-reduction","4-Way Split-KV Parallel Reduction",[14,2308,2309,2310,2312,2313,2316,2317,2324],{},"Borrowing the core concept from ",[48,2311,50],{},"'s ",[48,2314,2315],{},"fattn-vec.cuh",", we redesigned decode attention into a cooperative parallel reduction kernel (",[24,2318,2321],{"href":2319,"rel":2320},"https:\u002F\u002Fgithub.com\u002Fmihailtd\u002Fstrata\u002Fblob\u002Fmain\u002Fapps\u002Fruntime-next\u002Fsrc\u002Fkernels\u002Fattention_decode_split.hip",[28],[48,2322,2323],{},"attention_decode_split.hip","):",[307,2326,2327,2353,2773,2779],{},[300,2328,2329,2335,2336],{},[21,2330,2331,2332,101],{},"4 Cooperative Workers (",[48,2333,2334],{},"kv_split = 4",": Instead of 1 thread per dimension, 4 parallel worker threads collaborate on each output dimension:\n",[923,2337,2341],{"className":2338,"code":2339,"language":2340,"meta":931,"style":931},"language-c shiki shiki-themes github-light github-dark","int d     = tid % head_dim;\nint split = tid \u002F head_dim; \u002F\u002F 0..3\n","c",[48,2342,2343,2348],{"__ignoreMap":931},[469,2344,2345],{"class":1779,"line":1780},[469,2346,2347],{},"int d     = tid % head_dim;\n",[469,2349,2350],{"class":1779,"line":1786},[469,2351,2352],{},"int split = tid \u002F head_dim; \u002F\u002F 0..3\n",[300,2354,2355,2358,2359,2468,2469],{},[21,2356,2357],{},"Strided Segment Scanning",": Each worker thread scans exactly ",[469,2360,2362,2381],{"className":2361},[472],[469,2363,2365],{"className":2364},[476],[478,2366,2367],{"xmlns":480},[482,2368,2369,2378],{},[485,2370,2371],{},[2372,2373,2374,2376],"mfrac",{},[592,2375,838],{},[592,2377,1461],{},[502,2379,2380],{"encoding":504},"\\frac{1}{4}",[469,2382,2384],{"className":2383,"ariaHidden":510},[509],[469,2385,2387,2391],{"className":2386},[514],[469,2388],{"className":2389,"style":2390},[518],"height:1.1901em;vertical-align:-0.345em;",[469,2392,2394,2398,2465],{"className":2393},[523],[469,2395],{"className":2396},[529,2397],"nulldelimiter",[469,2399,2401],{"className":2400},[2372],[469,2402,2404,2456],{"className":2403},[632,633],[469,2405,2407,2453],{"className":2406},[637],[469,2408,2411,2427,2438],{"className":2409,"style":2410},[641],"height:0.8451em;",[469,2412,2414,2418],{"style":2413},"top:-2.655em;",[469,2415],{"className":2416,"style":2417},[649],"height:3em;",[469,2419,2421],{"className":2420},[654,655,656,657],[469,2422,2424],{"className":2423},[523,657],[469,2425,1461],{"className":2426},[523,657],[469,2428,2430,2433],{"style":2429},"top:-3.23em;",[469,2431],{"className":2432,"style":2417},[649],[469,2434],{"className":2435,"style":2437},[2436],"frac-line","border-bottom-width:0.04em;",[469,2439,2441,2444],{"style":2440},"top:-3.394em;",[469,2442],{"className":2443,"style":2417},[649],[469,2445,2447],{"className":2446},[654,655,656,657],[469,2448,2450],{"className":2449},[523,657],[469,2451,838],{"className":2452},[523,657],[469,2454,665],{"className":2455},[664],[469,2457,2459],{"className":2458},[637],[469,2460,2463],{"className":2461,"style":2462},[641],"height:0.345em;",[469,2464],{},[469,2466],{"className":2467},[537,2397]," of the KV cache length concurrently:\n",[469,2470,2472,2564],{"className":2471},[472],[469,2473,2475],{"className":2474},[476],[478,2476,2477],{"xmlns":480},[482,2478,2479,2561],{},[485,2480,2481,2484,2487,2490,2492,2495,2498,2500,2536,2539,2541,2543,2545,2548,2551,2553,2555,2557,2559],{},[1558,2482,2483],{},"partial",[492,2485,2486],{"stretchy":494},"[",[1558,2488,2489],{},"split",[492,2491,1563],{"separator":510},[488,2493,2494],{},"d",[492,2496,2497],{"stretchy":494},"]",[492,2499,1040],{},[2501,2502,2503,2506,2522],"msubsup",{},[492,2504,2505],{},"∑",[485,2507,2508,2511,2513,2515,2517,2520],{},[488,2509,2510],{},"j",[492,2512,1040],{},[1558,2514,2489],{},[492,2516,1563],{"separator":510},[1558,2518,2519],{},"step ",[592,2521,1461],{},[485,2523,2524,2526,2528,2530,2532,2534],{},[488,2525,1456],{},[488,2527,2234],{},[488,2529,2238],{"mathvariant":2237},[488,2531,2241],{},[488,2533,2244],{},[488,2535,2247],{},[1558,2537,2538],{},"prob",[492,2540,2486],{"stretchy":494},[488,2542,2510],{},[492,2544,2497],{"stretchy":494},[492,2546,2547],{},"⋅",[488,2549,2550],{},"V",[492,2552,2486],{"stretchy":494},[488,2554,2510],{},[492,2556,1563],{"separator":510},[488,2558,2494],{},[492,2560,2497],{"stretchy":494},[502,2562,2563],{"encoding":504},"\\text{partial}[\\text{split}, d] = \\sum_{j=\\text{split}, \\text{step } 4}^{kv\\_len} \\text{prob}[j] \\cdot V[j, d]",[469,2565,2567,2609,2746],{"className":2566,"ariaHidden":510},[509],[469,2568,2570,2573,2579,2582,2588,2591,2594,2597,2600,2603,2606],{"className":2569},[514],[469,2571],{"className":2572,"style":519},[518],[469,2574,2576],{"className":2575},[523,928],[469,2577,2483],{"className":2578},[523],[469,2580,2486],{"className":2581},[529],[469,2583,2585],{"className":2584},[523,928],[469,2586,2489],{"className":2587},[523],[469,2589,1563],{"className":2590},[1593],[469,2592],{"className":2593,"style":1597},[678],[469,2595,2494],{"className":2596},[523,524],[469,2598,2497],{"className":2599},[537],[469,2601],{"className":2602,"style":679},[678],[469,2604,1040],{"className":2605},[683],[469,2607],{"className":2608,"style":679},[678],[469,2610,2612,2616,2718,2721,2727,2730,2733,2736,2740,2743],{"className":2611},[514],[469,2613],{"className":2614,"style":2615},[518],"height:1.4853em;vertical-align:-0.4374em;",[469,2617,2620,2626],{"className":2618},[2619],"mop",[469,2621,2505],{"className":2622,"style":2625},[2619,2623,2624],"op-symbol","small-op","position:relative;top:0em;",[469,2627,2629],{"className":2628},[628],[469,2630,2632,2709],{"className":2631},[632,633],[469,2633,2635,2706],{"className":2634},[637],[469,2636,2639,2676],{"className":2637,"style":2638},[641],"height:1.0479em;",[469,2640,2642,2645],{"style":2641},"top:-2.3987em;margin-left:0em;margin-right:0.05em;",[469,2643],{"className":2644,"style":650},[649],[469,2646,2648],{"className":2647},[654,655,656,657],[469,2649,2651,2655,2658,2664,2667,2673],{"className":2650},[523,657],[469,2652,2510],{"className":2653,"style":2654},[523,524,657],"margin-right:0.0572em;",[469,2656,1040],{"className":2657},[683,657],[469,2659,2661],{"className":2660},[523,928,657],[469,2662,2489],{"className":2663},[523,657],[469,2665,1563],{"className":2666},[1593,657],[469,2668,2670],{"className":2669},[523,928,657],[469,2671,2519],{"className":2672},[523,657],[469,2674,1461],{"className":2675},[523,657],[469,2677,2679,2682],{"style":2678},"top:-3.2618em;margin-right:0.05em;",[469,2680],{"className":2681,"style":650},[649],[469,2683,2685],{"className":2684},[654,655,656,657],[469,2686,2688,2691,2694,2697,2700,2703],{"className":2687},[523,657],[469,2689,1456],{"className":2690,"style":1478},[523,524,657],[469,2692,2234],{"className":2693,"style":2275},[523,524,657],[469,2695,2238],{"className":2696,"style":525},[523,657],[469,2698,2241],{"className":2699,"style":2282},[523,524,657],[469,2701,2244],{"className":2702},[523,524,657],[469,2704,2247],{"className":2705},[523,524,657],[469,2707,665],{"className":2708},[664],[469,2710,2712],{"className":2711},[637],[469,2713,2716],{"className":2714,"style":2715},[641],"height:0.4374em;",[469,2717],{},[469,2719],{"className":2720,"style":1597},[678],[469,2722,2724],{"className":2723},[523,928],[469,2725,2538],{"className":2726},[523],[469,2728,2486],{"className":2729},[529],[469,2731,2510],{"className":2732,"style":2654},[523,524],[469,2734,2497],{"className":2735},[537],[469,2737],{"className":2738,"style":2739},[678],"margin-right:0.2222em;",[469,2741,2547],{"className":2742},[731],[469,2744],{"className":2745,"style":2739},[678],[469,2747,2749,2752,2755,2758,2761,2764,2767,2770],{"className":2748},[514],[469,2750],{"className":2751,"style":519},[518],[469,2753,2550],{"className":2754,"style":2739},[523,524],[469,2756,2486],{"className":2757},[529],[469,2759,2510],{"className":2760,"style":2654},[523,524],[469,2762,1563],{"className":2763},[1593],[469,2765],{"className":2766,"style":1597},[678],[469,2768,2494],{"className":2769},[523,524],[469,2771,2497],{"className":2772},[537],[300,2774,2775,2778],{},[21,2776,2777],{},"Warp Tree Reduction in LDS",": A workgroup tree reduction combines partial sums in GPU shared memory (Local Data Share, LDS).",[300,2780,2781,2784,2785,2788,2789,89],{},[21,2782,2783],{},"RDNA3 Hardware Bound",": ",[48,2786,2787],{},"head_dim (256) × kv_split (4) = 1,024 threads per block",". This lands exactly on the ",[21,2790,2791],{},"physical thread limit per compute block on AMD RDNA3 silicon",[14,2793,2794],{},[40,2795],{"alt":2796,"src":2797,"height":2798,"width":902},"Serial single-thread KV scan compared with four strided workers merged by a tree reduction in LDS | wide","\u002Fimages\u002Fblog\u002Fstrata-split-kv.svg#wide",620,[923,2800,2802],{"className":2338,"code":2801,"language":2340,"meta":931,"style":931},"\u002F\u002F Excerpt from apps\u002Fruntime-next\u002Fsrc\u002Fkernels\u002Fattention_decode_split.hip on GitHub\nextern \"C\" __global__ void attention_decode_split_bf16_kernel(\n    const unsigned short* q,\n    const unsigned short* k,\n    const unsigned short* v,\n    unsigned short* out,\n    int num_q_heads, int num_kv_heads,\n    const int* __restrict__ position_ptr,\n    int kv_stride, int head_dim, int kv_split, float scaling,\n    long long cache_stride, long long q_stride\n) {\n    \u002F\u002F 1024 threads per workgroup cooperating on reduction\n    int tid = threadIdx.x;\n    int d = tid % head_dim;\n    int split = tid \u002F head_dim;\n    \u002F\u002F ... parallel scan over kv_len \u002F kv_split ...\n    \u002F\u002F ... LDS shared memory warp combine ...\n}\n",[48,2803,2804,2809,2814,2819,2824,2829,2834,2839,2844,2849,2854,2860,2866,2872,2878,2884,2890,2896],{"__ignoreMap":931},[469,2805,2806],{"class":1779,"line":1780},[469,2807,2808],{},"\u002F\u002F Excerpt from apps\u002Fruntime-next\u002Fsrc\u002Fkernels\u002Fattention_decode_split.hip on GitHub\n",[469,2810,2811],{"class":1779,"line":1786},[469,2812,2813],{},"extern \"C\" __global__ void attention_decode_split_bf16_kernel(\n",[469,2815,2816],{"class":1779,"line":1792},[469,2817,2818],{},"    const unsigned short* q,\n",[469,2820,2821],{"class":1779,"line":1798},[469,2822,2823],{},"    const unsigned short* k,\n",[469,2825,2826],{"class":1779,"line":1804},[469,2827,2828],{},"    const unsigned short* v,\n",[469,2830,2831],{"class":1779,"line":1810},[469,2832,2833],{},"    unsigned short* out,\n",[469,2835,2836],{"class":1779,"line":1922},[469,2837,2838],{},"    int num_q_heads, int num_kv_heads,\n",[469,2840,2841],{"class":1779,"line":1928},[469,2842,2843],{},"    const int* __restrict__ position_ptr,\n",[469,2845,2846],{"class":1779,"line":1934},[469,2847,2848],{},"    int kv_stride, int head_dim, int kv_split, float scaling,\n",[469,2850,2851],{"class":1779,"line":1940},[469,2852,2853],{},"    long long cache_stride, long long q_stride\n",[469,2855,2857],{"class":1779,"line":2856},11,[469,2858,2859],{},") {\n",[469,2861,2863],{"class":1779,"line":2862},12,[469,2864,2865],{},"    \u002F\u002F 1024 threads per workgroup cooperating on reduction\n",[469,2867,2869],{"class":1779,"line":2868},13,[469,2870,2871],{},"    int tid = threadIdx.x;\n",[469,2873,2875],{"class":1779,"line":2874},14,[469,2876,2877],{},"    int d = tid % head_dim;\n",[469,2879,2881],{"class":1779,"line":2880},15,[469,2882,2883],{},"    int split = tid \u002F head_dim;\n",[469,2885,2887],{"class":1779,"line":2886},16,[469,2888,2889],{},"    \u002F\u002F ... parallel scan over kv_len \u002F kv_split ...\n",[469,2891,2893],{"class":1779,"line":2892},17,[469,2894,2895],{},"    \u002F\u002F ... LDS shared memory warp combine ...\n",[469,2897,2899],{"class":1779,"line":2898},18,[469,2900,1813],{},[14,2902,2903],{},"The microbenchmark measurements validated the redesign:",[114,2905,2906,2950],{},[117,2907,2908],{},[120,2909,2910,2941,2944,2947],{},[123,2911,2912,2913,101],{"align":129},"Cache Length (",[469,2914,2916,2929],{"className":2915},[472],[469,2917,2919],{"className":2918},[476],[478,2920,2921],{"xmlns":480},[482,2922,2923,2927],{},[485,2924,2925],{},[488,2926,498],{},[502,2928,498],{"encoding":504},[469,2930,2932],{"className":2931,"ariaHidden":510},[509],[469,2933,2935,2938],{"className":2934},[514],[469,2936],{"className":2937,"style":1055},[518],[469,2939,498],{"className":2940,"style":533},[523,524],[123,2942,2943],{"align":129},"Scalar Kernel Execution",[123,2945,2946],{"align":129},"4-Way Split Kernel Execution",[123,2948,2949],{"align":129},"Kernel Speedup",[150,2951,2952,2972,2992,3012],{},[120,2953,2954,2959,2962,2967],{},[155,2955,2956],{"align":129},[21,2957,2958],{},"64 tokens",[155,2960,2961],{"align":129},"14.34 μs",[155,2963,2964],{"align":129},[21,2965,2966],{},"8.68 μs",[155,2968,2969],{"align":129},[21,2970,2971],{},"1.65×",[120,2973,2974,2979,2982,2987],{},[155,2975,2976],{"align":129},[21,2977,2978],{},"128 tokens",[155,2980,2981],{"align":129},"19.55 μs",[155,2983,2984],{"align":129},[21,2985,2986],{},"10.37 μs",[155,2988,2989],{"align":129},[21,2990,2991],{},"1.88×",[120,2993,2994,2999,3002,3007],{},[155,2995,2996],{"align":129},[21,2997,2998],{},"256 tokens",[155,3000,3001],{"align":129},"32.46 μs",[155,3003,3004],{"align":129},[21,3005,3006],{},"14.04 μs",[155,3008,3009],{"align":129},[21,3010,3011],{},"2.31×",[120,3013,3014,3019,3022,3027],{},[155,3015,3016],{"align":129},[21,3017,3018],{},"354 tokens",[155,3020,3021],{"align":129},"44.81 μs",[155,3023,3024],{"align":129},[21,3025,3026],{},"16.79 μs",[155,3028,3029],{"align":129},[21,3030,3031],{},"2.67×",[14,3033,3034],{},[40,3035],{"alt":3036,"src":3037,"height":2798,"width":902},"Attention kernel time by cache length: scalar versus 4-way split, with speedups from 1.65x to 2.67x | wide","\u002Fimages\u002Fblog\u002Fstrata-kernel-speedup.svg#wide",[14,3039,3040],{},"When tested in the live HTTP server over full 350-token generation trajectories, decode throughput stayed flat:",[923,3042,3045],{"className":3043,"code":3044,"language":928},[926],"Segment Tokens 54  → 104:  83.05 tok\u002Fs\nSegment Tokens 104 → 154:  83.86 tok\u002Fs\nSegment Tokens 154 → 204:  83.80 tok\u002Fs\nSegment Tokens 204 → 254:  83.66 tok\u002Fs\nSegment Tokens 254 → 304:  83.41 tok\u002Fs\nSegment Tokens 304 → 354:  83.21 tok\u002Fs\n",[48,3046,3044],{"__ignoreMap":931},[14,3048,3049,3050,3053],{},"With 4-way split-KV attention active, the generation cadence held steady at ",[21,3051,3052],{},"~83.5 tok\u002Fs"," from token 1 to token 350.",[14,3055,3056],{},[40,3057],{"alt":3058,"src":3059,"height":3060,"width":902},"Decode tok\u002Fs across 350 tokens: scalar attention decays to 77.6, split-KV attention stays near 83.5 | wide","\u002Fimages\u002Fblog\u002Fstrata-decode-vs-depth.svg#wide",640,[286,3062],{},[289,3064,3066],{"id":3065},"act-5-auditing-the-baselines-the-speculative-decoding-trap","Act 5: Auditing the Baselines & The Speculative Decoding Trap",[14,3068,3069,3070,3072,3073,1221],{},"To guarantee that our lead over ",[48,3071,50],{}," was legitimate, we conducted an exhaustive configuration audit of ",[48,3074,100],{},[114,3076,3077,3090],{},[117,3078,3079],{},[120,3080,3081,3084,3087],{},[123,3082,3083],{"align":125},"Flag \u002F Parameter",[123,3085,3086],{"align":125},"Status in Baseline",[123,3088,3089],{"align":125},"Impact on Fairness",[150,3091,3092,3105,3117,3135,3151,3163,3175],{},[120,3093,3094,3099,3102],{},[155,3095,3096],{"align":125},[48,3097,3098],{},"-ngl 99",[155,3100,3101],{"align":125},"ACTIVE",[155,3103,3104],{"align":125},"Parity: 100% of all 32 layers pinned to VRAM. Zero CPU fallback.",[120,3106,3107,3112,3114],{},[155,3108,3109],{"align":125},[48,3110,3111],{},"-fa on",[155,3113,3101],{"align":125},[155,3115,3116],{"align":125},"Parity: Flash Attention enabled for full attention layers.",[120,3118,3119,3124,3126],{},[155,3120,3121],{"align":125},[48,3122,3123],{},"-ctk \u002F -ctv q8_0",[155,3125,3101],{"align":125},[155,3127,3128,3131,3132,3134],{"align":125},[21,3129,3130],{},"Advantage llama.cpp",": 8-bit quantized KV cache reduces bandwidth vs ",[21,3133,29],{},"'s unquantized BF16 buffers.",[120,3136,3137,3142,3144],{},[155,3138,3139],{"align":125},[48,3140,3141],{},"GGML_HIP_GRAPHS",[155,3143,3101],{"align":125},[155,3145,3146,3147,3150],{"align":125},"Parity: Verified compiled ON in ",[48,3148,3149],{},"CMakeCache.txt",". HIP graph replay active on gfx1100.",[120,3152,3153,3158,3160],{},[155,3154,3155],{"align":125},[48,3156,3157],{},"-b 512 \u002F -ub 512",[155,3159,3101],{"align":125},[155,3161,3162],{"align":125},"Parity: Matched batch and microbatch sizes.",[120,3164,3165,3170,3172],{},[155,3166,3167],{"align":125},[48,3168,3169],{},"-t 8",[155,3171,3101],{"align":125},[155,3173,3174],{"align":125},"Parity: 8 host CPU worker threads allocated.",[120,3176,3177,3182,3185],{},[155,3178,3179],{"align":125},[48,3180,3181],{},"--spec-type ngram-mod",[155,3183,3184],{"align":125},"TESTED SEPARATELY",[155,3186,3187],{"align":125},"Model-free speculative decoding evaluated below.",[14,3189,3190,3191,3193,3194,3197,3198,3200],{},"I enabled every known acceleration flag for ",[48,3192,50],{},". In fact, running with ",[48,3195,3196],{},"-ctk q8_0 -ctv q8_0"," gave ",[48,3199,50],{}," an inherent memory bandwidth advantage, as its attention layers read half as many KV bytes per token.",[940,3202,3204],{"id":3203},"why-speculative-decoding-failed-on-code-generation","Why Speculative Decoding Failed on Code Generation",[14,3206,3207,3208,3210,3211,2324],{},"Could ",[48,3209,50],{}," close the gap by enabling speculative decoding? We evaluated model-free n-gram speculation (",[48,3212,3181],{},[114,3214,3215,3231],{},[117,3216,3217],{},[120,3218,3219,3222,3225,3228],{},[123,3220,3221],{"align":125},"Configuration",[123,3223,3224],{"align":129},"Sustained tok\u002Fs",[123,3226,3227],{"align":129},"TTFT (ms)",[123,3229,3230],{"align":129},"Delta Throughput",[150,3232,3233,3255,3275],{},[120,3234,3235,3242,3247,3252],{},[155,3236,3237],{"align":125},[21,3238,3239,3241],{},[48,3240,50],{}," Baseline (No Speculation)",[155,3243,3244],{"align":129},[21,3245,3246],{},"72.73 tok\u002Fs",[155,3248,3249],{"align":129},[21,3250,3251],{},"108.0 ms",[155,3253,3254],{"align":129},"Baseline",[120,3256,3257,3266,3269,3272],{},[155,3258,3259,3261,3262,3265],{"align":125},[48,3260,50],{}," + ",[48,3263,3264],{},"ngram-mod"," (Run 1)",[155,3267,3268],{"align":129},"72.75 tok\u002Fs",[155,3270,3271],{"align":129},"107.4 ms",[155,3273,3274],{"align":129},"+0.02 tok\u002Fs (0.0%)",[120,3276,3277,3284,3287,3290],{},[155,3278,3279,3261,3281,3283],{"align":125},[48,3280,50],{},[48,3282,3264],{}," (Run 2)",[155,3285,3286],{"align":129},"71.04 tok\u002Fs",[155,3288,3289],{"align":129},"140.7 ms",[155,3291,3292],{"align":129},[21,3293,3294],{},"-1.69 tok\u002Fs (-2.3% degradation)",[14,3296,3297,3298,3300],{},"Telemetry logs emitted by ",[48,3299,100],{}," revealed the failure mechanism:",[923,3302,3306],{"className":3303,"code":3304,"language":3305,"meta":931,"style":931},"language-json shiki shiki-themes github-light github-dark","\"timings\": {\n    \"prompt_n\": 54,\n    \"predicted_n\": 350,\n    \"draft_n\": 64,\n    \"draft_n_accepted\": 5\n}\n","json",[48,3307,3308,3318,3332,3344,3356,3366],{"__ignoreMap":931},[469,3309,3310,3314],{"class":1779,"line":1780},[469,3311,3313],{"class":3312},"sZZnC","\"timings\"",[469,3315,3317],{"class":3316},"sVt8B",": {\n",[469,3319,3320,3324,3326,3329],{"class":1779,"line":1786},[469,3321,3323],{"class":3322},"sj4cs","    \"prompt_n\"",[469,3325,2784],{"class":3316},[469,3327,3328],{"class":3322},"54",[469,3330,3331],{"class":3316},",\n",[469,3333,3334,3337,3339,3342],{"class":1779,"line":1792},[469,3335,3336],{"class":3322},"    \"predicted_n\"",[469,3338,2784],{"class":3316},[469,3340,3341],{"class":3322},"350",[469,3343,3331],{"class":3316},[469,3345,3346,3349,3351,3354],{"class":1779,"line":1798},[469,3347,3348],{"class":3322},"    \"draft_n\"",[469,3350,2784],{"class":3316},[469,3352,3353],{"class":3322},"64",[469,3355,3331],{"class":3316},[469,3357,3358,3361,3363],{"class":1779,"line":1804},[469,3359,3360],{"class":3322},"    \"draft_n_accepted\"",[469,3362,2784],{"class":3316},[469,3364,3365],{"class":3322},"5\n",[469,3367,3368],{"class":1779,"line":1810},[469,3369,1813],{"class":3316},[14,3371,3372,3373,3376,3377,89],{},"The server accepted ",[21,3374,3375],{},"only 5"," of 64 drafted tokens, a ",[21,3378,3379],{},"7.8% acceptance rate",[14,3381,3382],{},[40,3383],{"alt":3384,"src":3385,"height":2798,"width":902},"64 drafted tokens with 5 accepted, and sustained decode tok\u002Fs with and without speculation | wide","\u002Fimages\u002Fblog\u002Fstrata-spec-decoding.svg#wide",[881,3387,3388],{},[14,3389,3390,3392,3395],{},[469,3391,887],{},[21,3393,3394],{},"The Guessing Assistant",":\nAn assistant attempts to guess the end of your sentence. If they shout out 64 words and 59 are wrong, you spend far more time stopping, correcting them, and restarting than if you had simply spoken at your normal pace.",[14,3397,3398,3399,3402],{},"In programming tasks (such as writing SQL migrations, FastAPI routes, and DuckDB analytics), tokens require exact syntactic and logical precision. Because every rejected token incurs verification passes and KV rollback overhead, an acceptance rate below 10% actually ",[21,3400,3401],{},"reduces"," throughput on an already-optimized decode loop.",[286,3404],{},[289,3406,3408],{"id":3407},"act-6-multi-size-scaling-08b-to-9b-the-memory-bandwidth-wall","Act 6: Multi-Size Scaling (0.8B to 9B) & The Memory Bandwidth Wall",[14,3410,3411,3412,89],{},"Does this speedup hold as models scale? We ran the identical head-to-head benchmark suite across the entire Qwen 3.5 dense family: ",[21,3413,3414],{},"0.8B, 2B, 4B, and 9B",[14,3416,3417],{},"All three engines loaded byte-identical BF16 weights, and all outputs passed the exactness gate against Hugging Face references:",[114,3419,3420,3451],{},[117,3421,3422],{},[120,3423,3424,3427,3431,3435,3439,3442,3445],{},[123,3425,3426],{"align":129},"Model Tier",[123,3428,3429,133],{"align":129},[21,3430,29],{},[123,3432,3433],{"align":129},[48,3434,50],{},[123,3436,3437],{"align":129},[48,3438,54],{},[123,3440,3441],{"align":129},"Speedup vs llama.cpp",[123,3443,3444],{"align":129},"Speedup vs Ollama",[123,3446,3447,3448,3450],{"align":129},"TTFT (",[21,3449,29],{}," vs llama.cpp)",[150,3452,3453,3486,3519,3549],{},[120,3454,3455,3460,3465,3468,3471,3476,3481],{},[155,3456,3457],{"align":129},[21,3458,3459],{},"0.8B",[155,3461,3462],{"align":129},[21,3463,3464],{},"285.9 tok\u002Fs",[155,3466,3467],{"align":129},"210.6 tok\u002Fs",[155,3469,3470],{"align":129},"222.5 tok\u002Fs",[155,3472,3473],{"align":129},[21,3474,3475],{},"1.36× (+35.7%)",[155,3477,3478],{"align":129},[21,3479,3480],{},"1.28× (+28.5%)",[155,3482,3483],{"align":129},[21,3484,3485],{},"17.0 ms vs 37.4 ms (-54.5%)",[120,3487,3488,3493,3498,3501,3504,3509,3514],{},[155,3489,3490],{"align":129},[21,3491,3492],{},"2B",[155,3494,3495],{"align":129},[21,3496,3497],{},"166.9 tok\u002Fs",[155,3499,3500],{"align":129},"136.4 tok\u002Fs",[155,3502,3503],{"align":129},"135.0 tok\u002Fs",[155,3505,3506],{"align":129},[21,3507,3508],{},"1.22× (+22.4%)",[155,3510,3511],{"align":129},[21,3512,3513],{},"1.24× (+23.7%)",[155,3515,3516],{"align":129},[21,3517,3518],{},"26.8 ms vs 54.2 ms (-50.6%)",[120,3520,3521,3526,3530,3532,3534,3539,3544],{},[155,3522,3523],{"align":129},[21,3524,3525],{},"4B",[155,3527,3528],{"align":129},[21,3529,164],{},[155,3531,167],{"align":129},[155,3533,170],{"align":129},[155,3535,3536],{"align":129},[21,3537,3538],{},"1.16× (+16.4%)",[155,3540,3541],{"align":129},[21,3542,3543],{},"1.15× (+15.3%)",[155,3545,3546],{"align":129},[21,3547,3548],{},"48.5 ms vs 121.0 ms (-59.9%)",[120,3550,3551,3556,3561,3564,3567,3572,3577],{},[155,3552,3553],{"align":129},[21,3554,3555],{},"9B",[155,3557,3558],{"align":129},[21,3559,3560],{},"49.1 tok\u002Fs",[155,3562,3563],{"align":129},"44.6 tok\u002Fs",[155,3565,3566],{"align":129},"45.1 tok\u002Fs",[155,3568,3569],{"align":129},[21,3570,3571],{},"1.10× (+10.2%)",[155,3573,3574],{"align":129},[21,3575,3576],{},"1.09× (+8.9%)",[155,3578,3579],{"align":129},[21,3580,3581],{},"81.0 ms vs 176.9 ms (-54.2%)",[14,3583,3584,3585,3587],{},"Across every size tier, ",[21,3586,29],{}," achieved the highest decode throughput and the lowest Time-To-First-Token.",[14,3589,3590],{},[40,3591],{"alt":3592,"src":3593,"height":3594,"width":902},"Decode tok\u002Fs for Strata, llama.cpp and Ollama at 0.8B, 2B, 4B and 9B, with Strata speedup shrinking from 35.7% to 10.2% | wide","\u002Fimages\u002Fblog\u002Fstrata-multisize.svg#wide",740,[14,3596,3597,3598,3601],{},"The trend is clear: ",[21,3599,3600],{},"the throughput speedup margin compresses as model size grows"," (from +35.7% at 0.8B down to +10.2% at 9B).",[940,3603,3605],{"id":3604},"the-kernel-execution-profile","The Kernel Execution Profile",[14,3607,3608,3609,1221],{},"To understand why the margin compresses, we profiled decode kernel execution across model sizes using ",[48,3610,1172],{},[923,3612,3615],{"className":3613,"code":3614,"language":928},[926],"Model Tier    Weight GEMV Kernels    GDN Recurrent Update    Other Ops (Attn, Norms, RoPE)\n0.8B                72.66%                 13.59%                       13.75%\n4B                  88.85%                  5.40%                        5.75%\n9B                  93.64%                  3.02%                        3.34%\n",[48,3616,3614],{"__ignoreMap":931},[14,3618,3619],{},"Two architectural realities explain this compression:",[297,3621,3622,3980],{},[300,3623,3624,3627,3628,97,3688,3769,3770,557,3801,3881,3882,3976,3977,89],{},[21,3625,3626],{},"Quadratic Parameter Scaling",":\nWeight matrix sizes scale with ",[469,3629,3631,3651],{"className":3630},[472],[469,3632,3634],{"className":3633},[476],[478,3635,3636],{"xmlns":480},[482,3637,3638,3648],{},[485,3639,3640,3643,3645],{},[1558,3641,3642],{},"hidden_size",[492,3644,597],{},[1558,3646,3647],{},"intermediate_size",[502,3649,3650],{"encoding":504},"\\text{hidden\\_size} \\times \\text{intermediate\\_size}",[469,3652,3654,3676],{"className":3653,"ariaHidden":510},[509],[469,3655,3657,3661,3667,3670,3673],{"className":3656},[514],[469,3658],{"className":3659,"style":3660},[518],"height:1.0044em;vertical-align:-0.31em;",[469,3662,3664],{"className":3663},[523,928],[469,3665,3642],{"className":3666},[523],[469,3668],{"className":3669,"style":2739},[678],[469,3671,597],{"className":3672},[731],[469,3674],{"className":3675,"style":2739},[678],[469,3677,3679,3682],{"className":3678},[514],[469,3680],{"className":3681,"style":3660},[518],[469,3683,3685],{"className":3684},[523,928],[469,3686,3647],{"className":3687},[523],[469,3689,3691,3720],{"className":3690},[472],[469,3692,3694],{"className":3693},[476],[478,3695,3696],{"xmlns":480},[482,3697,3698,3717],{},[485,3699,3700,3702,3704,3707,3709,3712,3714],{},[592,3701,838],{},[492,3703,1563],{"separator":510},[592,3705,3706],{},"024",[492,3708,597],{},[592,3710,3711],{},"3",[492,3713,1563],{"separator":510},[592,3715,3716],{},"584",[502,3718,3719],{"encoding":504},"1,024 \\times 3,584",[469,3721,3723,3751],{"className":3722,"ariaHidden":510},[509],[469,3724,3726,3730,3733,3736,3739,3742,3745,3748],{"className":3725},[514],[469,3727],{"className":3728,"style":3729},[518],"height:0.8389em;vertical-align:-0.1944em;",[469,3731,838],{"className":3732},[523],[469,3734,1563],{"className":3735},[1593],[469,3737],{"className":3738,"style":1597},[678],[469,3740,3706],{"className":3741},[523],[469,3743],{"className":3744,"style":2739},[678],[469,3746,597],{"className":3747},[731],[469,3749],{"className":3750,"style":2739},[678],[469,3752,3754,3757,3760,3763,3766],{"className":3753},[514],[469,3755],{"className":3756,"style":3729},[518],[469,3758,3711],{"className":3759},[523],[469,3761,1563],{"className":3762},[1593],[469,3764],{"className":3765,"style":1597},[678],[469,3767,3716],{"className":3768},[523]," at 0.8B ",[469,3771,3773,3788],{"className":3772},[472],[469,3774,3776],{"className":3775},[476],[478,3777,3778],{"xmlns":480},[482,3779,3780,3785],{},[485,3781,3782],{},[492,3783,3784],{},"→",[502,3786,3787],{"encoding":504},"\\to",[469,3789,3791],{"className":3790,"ariaHidden":510},[509],[469,3792,3794,3798],{"className":3793},[514],[469,3795],{"className":3796,"style":3797},[518],"height:0.3669em;",[469,3799,3784],{"className":3800},[683],[469,3802,3804,3833],{"className":3803},[472],[469,3805,3807],{"className":3806},[476],[478,3808,3809],{"xmlns":480},[482,3810,3811,3830],{},[485,3812,3813,3815,3817,3820,3822,3825,3827],{},[592,3814,1461],{},[492,3816,1563],{"separator":510},[592,3818,3819],{},"096",[492,3821,597],{},[592,3823,3824],{},"12",[492,3826,1563],{"separator":510},[592,3828,3829],{},"288",[502,3831,3832],{"encoding":504},"4,096 \\times 12,288",[469,3834,3836,3863],{"className":3835,"ariaHidden":510},[509],[469,3837,3839,3842,3845,3848,3851,3854,3857,3860],{"className":3838},[514],[469,3840],{"className":3841,"style":3729},[518],[469,3843,1461],{"className":3844},[523],[469,3846,1563],{"className":3847},[1593],[469,3849],{"className":3850,"style":1597},[678],[469,3852,3819],{"className":3853},[523],[469,3855],{"className":3856,"style":2739},[678],[469,3858,597],{"className":3859},[731],[469,3861],{"className":3862,"style":2739},[678],[469,3864,3866,3869,3872,3875,3878],{"className":3865},[514],[469,3867],{"className":3868,"style":3729},[518],[469,3870,3824],{"className":3871},[523],[469,3873,1563],{"className":3874},[1593],[469,3876],{"className":3877,"style":1597},[678],[469,3879,3829],{"className":3880},[523]," at 9B). By contrast, GDN recurrence cost scales with ",[469,3883,3885,3910],{"className":3884},[472],[469,3886,3888],{"className":3887},[476],[478,3889,3890],{"xmlns":480},[482,3891,3892,3907],{},[485,3893,3894,3897,3899],{},[1558,3895,3896],{},"heads",[492,3898,597],{},[583,3900,3901,3904],{},[1558,3902,3903],{},"head_dim",[592,3905,3906],{},"2",[502,3908,3909],{"encoding":504},"\\text{heads} \\times \\text{head\\_dim}^2",[469,3911,3913,3935],{"className":3912,"ariaHidden":510},[509],[469,3914,3916,3920,3926,3929,3932],{"className":3915},[514],[469,3917],{"className":3918,"style":3919},[518],"height:0.7778em;vertical-align:-0.0833em;",[469,3921,3923],{"className":3922},[523,928],[469,3924,3896],{"className":3925},[523],[469,3927],{"className":3928,"style":2739},[678],[469,3930,597],{"className":3931},[731],[469,3933],{"className":3934,"style":2739},[678],[469,3936,3938,3942],{"className":3937},[514],[469,3939],{"className":3940,"style":3941},[518],"height:1.2084em;vertical-align:-0.31em;",[469,3943,3945,3951],{"className":3944},[523],[469,3946,3948],{"className":3947},[523,928],[469,3949,3903],{"className":3950},[523],[469,3952,3954],{"className":3953},[628],[469,3955,3957],{"className":3956},[632],[469,3958,3960],{"className":3959},[637],[469,3961,3964],{"className":3962,"style":3963},[641],"height:0.8984em;",[469,3965,3967,3970],{"style":3966},"top:-3.1473em;margin-right:0.05em;",[469,3968],{"className":3969,"style":650},[649],[469,3971,3973],{"className":3972},[654,655,656,657],[469,3974,3906],{"className":3975},[523,657],", which stays constant. At 9B, GEMV consumes ",[21,3978,3979],{},"over 93.6% of total GPU execution time",[300,3981,3982,3985,3986,4036,4037,4173,4174,4330,4332,4333,4336,4337,89],{},[21,3983,3984],{},"The 960 GB\u002Fs Physical Memory Bandwidth Wall",":\nDuring single-token decode (",[469,3987,3989,4006],{"className":3988},[472],[469,3990,3992],{"className":3991},[476],[478,3993,3994],{"xmlns":480},[482,3995,3996,4004],{},[485,3997,3998,4000,4002],{},[488,3999,1037],{},[492,4001,1040],{},[592,4003,838],{},[502,4005,1045],{"encoding":504},[469,4007,4009,4027],{"className":4008,"ariaHidden":510},[509],[469,4010,4012,4015,4018,4021,4024],{"className":4011},[514],[469,4013],{"className":4014,"style":1055},[518],[469,4016,1037],{"className":4017,"style":1059},[523,524],[469,4019],{"className":4020,"style":679},[678],[469,4022,1040],{"className":4023},[683],[469,4025],{"className":4026,"style":679},[678],[469,4028,4030,4033],{"className":4029},[514],[469,4031],{"className":4032,"style":1075},[518],[469,4034,838],{"className":4035},[523],"), the GPU must stream every parameter in the model from VRAM into the compute units exactly once per generated token:\n",[469,4038,4040,4065],{"className":4039},[472],[469,4041,4043],{"className":4042},[476],[478,4044,4045],{"xmlns":480},[482,4046,4047,4062],{},[485,4048,4049,4052,4054],{},[1558,4050,4051],{},"Theoretical Max tok\u002Fs",[492,4053,1040],{},[2372,4055,4056,4059],{},[1558,4057,4058],{},"Memory Bandwidth (GB\u002Fs)",[1558,4060,4061],{},"Model Size in VRAM (GB)",[502,4063,4064],{"encoding":504},"\\text{Theoretical Max tok\u002Fs} = \\frac{\\text{Memory Bandwidth (GB\u002Fs)}}{\\text{Model Size in VRAM (GB)}}",[469,4066,4068,4089],{"className":4067,"ariaHidden":510},[509],[469,4069,4071,4074,4080,4083,4086],{"className":4070},[514],[469,4072],{"className":4073,"style":519},[518],[469,4075,4077],{"className":4076},[523,928],[469,4078,4051],{"className":4079},[523],[469,4081],{"className":4082,"style":679},[678],[469,4084,1040],{"className":4085},[683],[469,4087],{"className":4088,"style":679},[678],[469,4090,4092,4096],{"className":4091},[514],[469,4093],{"className":4094,"style":4095},[518],"height:1.53em;vertical-align:-0.52em;",[469,4097,4099,4102,4170],{"className":4098},[523],[469,4100],{"className":4101},[529,2397],[469,4103,4105],{"className":4104},[2372],[469,4106,4108,4161],{"className":4107},[632,633],[469,4109,4111,4158],{"className":4110},[637],[469,4112,4115,4132,4140],{"className":4113,"style":4114},[641],"height:1.01em;",[469,4116,4117,4120],{"style":2413},[469,4118],{"className":4119,"style":2417},[649],[469,4121,4123],{"className":4122},[654,655,656,657],[469,4124,4126],{"className":4125},[523,657],[469,4127,4129],{"className":4128},[523,928,657],[469,4130,4061],{"className":4131},[523,657],[469,4133,4134,4137],{"style":2429},[469,4135],{"className":4136,"style":2417},[649],[469,4138],{"className":4139,"style":2437},[2436],[469,4141,4143,4146],{"style":4142},"top:-3.485em;",[469,4144],{"className":4145,"style":2417},[649],[469,4147,4149],{"className":4148},[654,655,656,657],[469,4150,4152],{"className":4151},[523,657],[469,4153,4155],{"className":4154},[523,928,657],[469,4156,4058],{"className":4157},[523,657],[469,4159,665],{"className":4160},[664],[469,4162,4164],{"className":4163},[637],[469,4165,4168],{"className":4166,"style":4167},[641],"height:0.52em;",[469,4169],{},[469,4171],{"className":4172},[537,2397],"\nAt 9B in BF16 (~18.2 GB parameter buffer), the absolute theoretical ceiling on a 960 GB\u002Fs bus is:\n",[469,4175,4177,4216],{"className":4176},[472],[469,4178,4180],{"className":4179},[476],[478,4181,4182],{"xmlns":480},[482,4183,4184,4213],{},[485,4185,4186,4204,4207,4210],{},[2372,4187,4188,4196],{},[485,4189,4190,4193],{},[592,4191,4192],{},"960",[1558,4194,4195],{}," GB\u002Fs",[485,4197,4198,4201],{},[592,4199,4200],{},"18.2",[1558,4202,4203],{}," GB",[492,4205,4206],{},"≈",[592,4208,4209],{},"52.7",[1558,4211,4212],{}," tok\u002Fs",[502,4214,4215],{"encoding":504},"\\frac{960 \\text{ GB\u002Fs}}{18.2 \\text{ GB}} \\approx 52.7 \\text{ tok\u002Fs}",[469,4217,4219,4315],{"className":4218,"ariaHidden":510},[509],[469,4220,4222,4226,4306,4309,4312],{"className":4221},[514],[469,4223],{"className":4224,"style":4225},[518],"height:1.355em;vertical-align:-0.345em;",[469,4227,4229,4232,4303],{"className":4228},[523],[469,4230],{"className":4231},[529,2397],[469,4233,4235],{"className":4234},[2372],[469,4236,4238,4295],{"className":4237},[632,633],[469,4239,4241,4292],{"className":4240},[637],[469,4242,4244,4264,4272],{"className":4243,"style":4114},[641],[469,4245,4246,4249],{"style":2413},[469,4247],{"className":4248,"style":2417},[649],[469,4250,4252],{"className":4251},[654,655,656,657],[469,4253,4255,4258],{"className":4254},[523,657],[469,4256,4200],{"className":4257},[523,657],[469,4259,4261],{"className":4260},[523,928,657],[469,4262,4203],{"className":4263},[523,657],[469,4265,4266,4269],{"style":2429},[469,4267],{"className":4268,"style":2417},[649],[469,4270],{"className":4271,"style":2437},[2436],[469,4273,4274,4277],{"style":4142},[469,4275],{"className":4276,"style":2417},[649],[469,4278,4280],{"className":4279},[654,655,656,657],[469,4281,4283,4286],{"className":4282},[523,657],[469,4284,4192],{"className":4285},[523,657],[469,4287,4289],{"className":4288},[523,928,657],[469,4290,4195],{"className":4291},[523,657],[469,4293,665],{"className":4294},[664],[469,4296,4298],{"className":4297},[637],[469,4299,4301],{"className":4300,"style":2462},[641],[469,4302],{},[469,4304],{"className":4305},[537,2397],[469,4307],{"className":4308,"style":679},[678],[469,4310,4206],{"className":4311},[683],[469,4313],{"className":4314,"style":679},[678],[469,4316,4318,4321,4324],{"className":4317},[514],[469,4319],{"className":4320,"style":519},[518],[469,4322,4209],{"className":4323},[523],[469,4325,4327],{"className":4326},[523,928],[469,4328,4212],{"className":4329},[523],[21,4331,29],{}," achieved ",[21,4334,4335],{},"49.12 tok\u002Fs","—which represents ",[21,4338,4339],{},"93.2% of the theoretical physical memory bandwidth of the GPU",[14,4341,4342],{},[40,4343],{"alt":4344,"src":4345,"height":4346,"width":902},"Kernel time share by model size and 9B decode at 93.2% of the memory bandwidth ceiling | wide","\u002Fimages\u002Fblog\u002Fstrata-bandwidth-wall.svg#wide",720,[14,4348,4349,4350,51,4352,4354],{},"When an engine operates at 93% of physical hardware wire limits, there is almost no software headroom left to extract. At 9B, both ",[21,4351,29],{},[48,4353,50],{}," are memory-bandwidth bound.",[286,4356],{},[289,4358,4360],{"id":4359},"️-what-broke-hard-limits-caveats","⚠️ What Broke: Hard Limits & Caveats",[14,4362,4363],{},"An honest report documents what broke and where the hard ceilings sit. These are mine:",[940,4365,4367],{"id":4366},"_1-the-14080-token-shared-memory-lds-ceiling","1. The 14,080-Token Shared Memory (LDS) Ceiling",[14,4369,4370,4371,4376],{},"In our 4-way split attention kernel (",[24,4372,4374],{"href":2319,"rel":4373},[28],[48,4375,2323],{},"), shared memory is dynamically allocated across the workgroup:",[923,4378,4380],{"className":2338,"code":4379,"language":2340,"meta":931,"style":931},"size_t shmem_bytes = (head_dim + kv_stride + threads + head_dim * kv_split) * sizeof(float);\n",[48,4381,4382],{"__ignoreMap":931},[469,4383,4384],{"class":1779,"line":1780},[469,4385,4379],{},[14,4387,4388,4389,4391,4392,89],{},"On AMD RDNA3 (Navi 31, ",[48,4390,64],{},"), each Compute Unit workgroup has ",[21,4393,4394],{},"64 KB of Local Data Share (LDS)",[14,4396,4397,4398,51,4401,4403,4404],{},"With ",[48,4399,4400],{},"threads = 1024",[48,4402,2178],{},", the longest sequence stride that fits into 64 KB is:\n",[469,4405,4407,4474],{"className":4406},[472],[469,4408,4410],{"className":4409},[476],[478,4411,4412],{"xmlns":480},[482,4413,4414,4471],{},[485,4415,4416,4419,4422,4436,4439,4442,4444,4446,4448,4450,4452,4454,4456,4458,4460,4463,4465,4468],{},[1558,4417,4418],{},"max_seq_len",[492,4420,4421],{},"≤",[2372,4423,4424,4434],{},[485,4425,4426,4429,4431],{},[592,4427,4428],{},"65",[492,4430,1563],{"separator":510},[592,4432,4433],{},"536",[592,4435,1461],{},[492,4437,4438],{},"−",[592,4440,4441],{},"256",[492,4443,4438],{},[492,4445,495],{"stretchy":494},[592,4447,3906],{},[492,4449,597],{},[592,4451,4441],{},[492,4453,597],{},[592,4455,1461],{},[492,4457,101],{"stretchy":494},[492,4459,1040],{},[592,4461,4462],{},"14",[492,4464,1563],{"separator":510},[592,4466,4467],{},"080",[1558,4469,4470],{}," tokens",[502,4472,4473],{"encoding":504},"\\text{max\\_seq\\_len} \\le \\frac{65,536}{4} - 256 - (2 \\times 256 \\times 4) = 14,080 \\text{ tokens}",[469,4475,4477,4498,4590,4609,4630,4648,4669],{"className":4476,"ariaHidden":510},[509],[469,4478,4480,4483,4489,4492,4495],{"className":4479},[514],[469,4481],{"className":4482,"style":3660},[518],[469,4484,4486],{"className":4485},[523,928],[469,4487,4418],{"className":4488},[523],[469,4490],{"className":4491,"style":679},[678],[469,4493,4421],{"className":4494},[683],[469,4496],{"className":4497,"style":679},[678],[469,4499,4501,4505,4581,4584,4587],{"className":4500},[514],[469,4502],{"className":4503,"style":4504},[518],"height:1.2422em;vertical-align:-0.345em;",[469,4506,4508,4511,4578],{"className":4507},[523],[469,4509],{"className":4510},[529,2397],[469,4512,4514],{"className":4513},[2372],[469,4515,4517,4570],{"className":4516},[632,633],[469,4518,4520,4567],{"className":4519},[637],[469,4521,4524,4538,4546],{"className":4522,"style":4523},[641],"height:0.8972em;",[469,4525,4526,4529],{"style":2413},[469,4527],{"className":4528,"style":2417},[649],[469,4530,4532],{"className":4531},[654,655,656,657],[469,4533,4535],{"className":4534},[523,657],[469,4536,1461],{"className":4537},[523,657],[469,4539,4540,4543],{"style":2429},[469,4541],{"className":4542,"style":2417},[649],[469,4544],{"className":4545,"style":2437},[2436],[469,4547,4549,4552],{"style":4548},"top:-3.4461em;",[469,4550],{"className":4551,"style":2417},[649],[469,4553,4555],{"className":4554},[654,655,656,657],[469,4556,4558,4561,4564],{"className":4557},[523,657],[469,4559,4428],{"className":4560},[523,657],[469,4562,1563],{"className":4563},[1593,657],[469,4565,4433],{"className":4566},[523,657],[469,4568,665],{"className":4569},[664],[469,4571,4573],{"className":4572},[637],[469,4574,4576],{"className":4575,"style":2462},[641],[469,4577],{},[469,4579],{"className":4580},[537,2397],[469,4582],{"className":4583,"style":2739},[678],[469,4585,4438],{"className":4586},[731],[469,4588],{"className":4589,"style":2739},[678],[469,4591,4593,4597,4600,4603,4606],{"className":4592},[514],[469,4594],{"className":4595,"style":4596},[518],"height:0.7278em;vertical-align:-0.0833em;",[469,4598,4441],{"className":4599},[523],[469,4601],{"className":4602,"style":2739},[678],[469,4604,4438],{"className":4605},[731],[469,4607],{"className":4608,"style":2739},[678],[469,4610,4612,4615,4618,4621,4624,4627],{"className":4611},[514],[469,4613],{"className":4614,"style":519},[518],[469,4616,495],{"className":4617},[529],[469,4619,3906],{"className":4620},[523],[469,4622],{"className":4623,"style":2739},[678],[469,4625,597],{"className":4626},[731],[469,4628],{"className":4629,"style":2739},[678],[469,4631,4633,4636,4639,4642,4645],{"className":4632},[514],[469,4634],{"className":4635,"style":4596},[518],[469,4637,4441],{"className":4638},[523],[469,4640],{"className":4641,"style":2739},[678],[469,4643,597],{"className":4644},[731],[469,4646],{"className":4647,"style":2739},[678],[469,4649,4651,4654,4657,4660,4663,4666],{"className":4650},[514],[469,4652],{"className":4653,"style":519},[518],[469,4655,1461],{"className":4656},[523],[469,4658,101],{"className":4659},[537],[469,4661],{"className":4662,"style":679},[678],[469,4664,1040],{"className":4665},[683],[469,4667],{"className":4668,"style":679},[678],[469,4670,4672,4676,4679,4682,4685,4688],{"className":4671},[514],[469,4673],{"className":4674,"style":4675},[518],"height:0.8889em;vertical-align:-0.1944em;",[469,4677,4462],{"className":4678},[523],[469,4680,1563],{"className":4681},[1593],[469,4683],{"className":4684,"style":1597},[678],[469,4686,4467],{"className":4687},[523],[469,4689,4691],{"className":4690},[523,928],[469,4692,4470],{"className":4693},[523],[14,4695,4696,4697,89],{},"We discovered this the hard way: when we raised the context window to 32,768, prefill processed 32,000 tokens smoothly, and then the first decode step failed with a fatal HIP error: ",[48,4698,4699],{},"hipErrorInvalidValue",[14,4701,4702,4703,4705,4706,4709],{},"Context length in ",[21,4704,29],{}," is currently bounded at ",[21,4707,4708],{},"12,288–14,080 tokens",". Crossing this threshold requires refactoring decode attention into a multi-block Flash-Decode architecture that combines partial sums across separate workgroups.",[14,4711,4712],{},[40,4713],{"alt":4714,"src":4715,"height":4716,"width":902},"LDS budget for the split-KV kernel: 14,080 tokens fit in 64 KB, 32,768 tokens need 137 KB | wide","\u002Fimages\u002Fblog\u002Fstrata-lds-ceiling.svg#wide",590,[940,4718,4720],{"id":4719},"_2-guarding-against-gpu-page-faults","2. Guarding Against GPU Page Faults",[14,4722,4723,4724,4727,4728,4731],{},"In early prototypes, if an incoming HTTP request requested ",[48,4725,4726],{},"max_tokens: 4096"," against an existing 7,000-token prompt, decode stepped right past the pre-allocated KV buffer, triggering an ",[21,4729,4730],{},"uncorrectable GPU page fault"," that crashed the ROCm driver and required a host reboot.",[14,4733,4734],{},"We hardened the generation loop with an explicit safety boundary:",[923,4736,4738],{"className":1772,"code":4737,"language":1774,"meta":931,"style":931},"if engine.context_exhausted() {\n    chunk.finish_reason = Some(\"length\");\n    break;\n}\n",[48,4739,4740,4745,4750,4755],{"__ignoreMap":931},[469,4741,4742],{"class":1779,"line":1780},[469,4743,4744],{},"if engine.context_exhausted() {\n",[469,4746,4747],{"class":1779,"line":1786},[469,4748,4749],{},"    chunk.finish_reason = Some(\"length\");\n",[469,4751,4752],{"class":1779,"line":1792},[469,4753,4754],{},"    break;\n",[469,4756,4757],{"class":1779,"line":1798},[469,4758,1813],{},[14,4760,4761,4762,4765],{},"The server now cleanly halts generation with standard HTTP ",[48,4763,4764],{},"finish_reason: \"length\"",", preserving GPU memory integrity.",[940,4767,4769],{"id":4768},"_3-bf16-vs-quantization","3. BF16 vs Quantization",[14,4771,4772,4773,4776],{},"Every benchmark in this article used ",[21,4774,4775],{},"unquantized BF16 weights",". Unquantized inference is a pure test of memory bandwidth, kernel fusion, and host dispatch. Weight quantization (such as W4A16 INT4 or Q4_K_M) introduces on-the-fly dequantization math, integer unpacking overhead, and cache pressure—which presents a distinct set of tradeoffs (covered in Part 2).",[940,4778,4780],{"id":4779},"_4-single-sequence-vs-high-throughput-batched-serving","4. Single-Sequence vs High-Throughput Batched Serving",[14,4782,4783,4784,4786,4787,4843,4844,4895,4896,4947],{},"I built ",[21,4785,29],{}," specifically for ",[21,4788,4789,4790,101],{},"interactive, low-latency agentic streaming (batch size ",[469,4791,4793,4812],{"className":4792},[472],[469,4794,4796],{"className":4795},[476],[478,4797,4798],{"xmlns":480},[482,4799,4800,4809],{},[485,4801,4802,4805,4807],{},[488,4803,4804],{},"B",[492,4806,1040],{},[592,4808,838],{},[502,4810,4811],{"encoding":504},"B=1",[469,4813,4815,4834],{"className":4814,"ariaHidden":510},[509],[469,4816,4818,4821,4825,4828,4831],{"className":4817},[514],[469,4819],{"className":4820,"style":1055},[518],[469,4822,4804],{"className":4823,"style":4824},[523,524],"margin-right:0.0502em;",[469,4826],{"className":4827,"style":679},[678],[469,4829,1040],{"className":4830},[683],[469,4832],{"className":4833,"style":679},[678],[469,4835,4837,4840],{"className":4836},[514],[469,4838],{"className":4839,"style":1075},[518],[469,4841,838],{"className":4842},[523],". When scaling to heavy server batching (",[469,4845,4847,4865],{"className":4846},[472],[469,4848,4850],{"className":4849},[476],[478,4851,4852],{"xmlns":480},[482,4853,4854,4862],{},[485,4855,4856,4858,4860],{},[488,4857,4804],{},[492,4859,1040],{},[592,4861,594],{},[502,4863,4864],{"encoding":504},"B=32",[469,4866,4868,4886],{"className":4867,"ariaHidden":510},[509],[469,4869,4871,4874,4877,4880,4883],{"className":4870},[514],[469,4872],{"className":4873,"style":1055},[518],[469,4875,4804],{"className":4876,"style":4824},[523,524],[469,4878],{"className":4879,"style":679},[678],[469,4881,1040],{"className":4882},[683],[469,4884],{"className":4885,"style":679},[678],[469,4887,4889,4892],{"className":4888},[514],[469,4890],{"className":4891,"style":1075},[518],[469,4893,594],{"className":4894},[523]," or ",[469,4897,4899,4917],{"className":4898},[472],[469,4900,4902],{"className":4901},[476],[478,4903,4904],{"xmlns":480},[482,4905,4906,4914],{},[485,4907,4908,4910,4912],{},[488,4909,4804],{},[492,4911,1040],{},[592,4913,3353],{},[502,4915,4916],{"encoding":504},"B=64",[469,4918,4920,4938],{"className":4919,"ariaHidden":510},[509],[469,4921,4923,4926,4929,4932,4935],{"className":4922},[514],[469,4924],{"className":4925,"style":1055},[518],[469,4927,4804],{"className":4928,"style":4824},[523,524],[469,4930],{"className":4931,"style":679},[678],[469,4933,1040],{"className":4934},[683],[469,4936],{"className":4937,"style":679},[678],[469,4939,4941,4944],{"className":4940},[514],[469,4942],{"className":4943,"style":1075},[518],[469,4945,3353],{"className":4946},[523],"), execution shifts from memory-bandwidth bound to compute bound. While Strata supports batched decode, multi-tenant serving introduces different trade-offs in prefill scheduling and KV memory fragmentation.",[286,4949],{},[289,4951,4953],{"id":4952},"key-engineering-takeaways","🎯 Key Engineering Takeaways",[14,4955,4956,4957,4959],{},"Building ",[21,4958,29],{}," from scratch taught us four fundamental principles of GPU systems engineering and local AI architecture:",[297,4961,4962,4968,4974,4980],{},[300,4963,4964,4967],{},[21,4965,4966],{},"Host-Side Overhead Can Dwarf GPU Compute",":\nOur initial prefill implementation spent 50 ms in CPU driver queues issuing 3,300 small kernels, overshadowing the 20 ms of actual matrix compute. Batching host dispatch is just as critical as optimizing GEMM math.",[300,4969,4970,4973],{},[21,4971,4972],{},"Default Socket Buffering Silently Breaks Streaming",":\nAn innocent 8KB buffer inside an HTTP crate introduced a 540 ms latency penalty by silently hoarding the first 55 tokens. Streaming runtimes must take direct control of the TCP socket with immediate per-frame flushes.",[300,4975,4976,4979],{},[21,4977,4978],{},"Isolate First, Hypothesize Second",":\nWhen streaming throughput drifted from 90 to 79 tok\u002Fs, our intuition blamed per-token network flushing. An in-process A\u002FB\u002FC isolation test proved network framing accounted for less than 0.3% of runtime, directing our attention to the real culprit: the serial KV cache reduction loop.",[300,4981,4982,4985],{},[21,4983,4984],{},"Consumer AMD Silicon Is Genuinely Capable",":\nWhen programmed directly in native HIP and safe Rust, consumer AMD GPUs (like the Radeon RX 7900 XTX) deliver top-tier throughput and latency. You do not need enterprise datacenter silicon or closed proprietary APIs to achieve bleeding-edge local LLM inference performance.",[286,4987],{},[14,4989,4990],{},[17,4991,1754,4992,4995],{},[21,4993,4994],{},"Part 2",", we examine what happened when we took Strata into quantized territory: why our initial W4A16 INT4 kernel lost to llama.cpp by 22%, and how custom dequant-gemv kernels turned it into a clean win across all model sizes.",[286,4997],{},[289,4999,5001],{"id":5000},"about-the-author","👨‍💻 About the Author",[14,5003,5004,5007,5008,359,5011,365,5014,5017],{},[21,5005,5006],{},"Mihai Farcas"," is a ",[21,5009,5010],{},"Software Architect",[21,5012,5013],{},"AI Systems Engineer",[21,5015,5016],{},"GPU Kernel Developer"," specializing in high-performance distributed systems, bare-metal hardware acceleration (AMD ROCm\u002FHIP, NVIDIA CUDA), and local LLM infrastructure.",[307,5019,5020,5035,5045],{},[300,5021,5022,2784,5025,5030,5031,101],{},[21,5023,5024],{},"GitHub",[24,5026,5029],{"href":5027,"rel":5028},"https:\u002F\u002Fgithub.com\u002Fmihailtd",[28],"github.com\u002Fmihailtd"," (Explore the ",[24,5032,5034],{"href":26,"rel":5033},[28],"Strata Repository",[300,5036,5037,2784,5040],{},[21,5038,5039],{},"Engineering Blog & Portfolio",[24,5041,5044],{"href":5042,"rel":5043},"https:\u002F\u002Fmihai.ltd",[28],"mihai.ltd",[300,5046,5047,5050],{},[21,5048,5049],{},"Domain Focus",": Local AI Architecture, High-Throughput Inference Engines, Custom GPU Kernels, and Zero-Allocation Systems Programming in Rust and C++.",[5052,5053,5054],"style",{},"html .default .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .shiki span {color: var(--shiki-default);background: var(--shiki-default-bg);font-style: var(--shiki-default-font-style);font-weight: var(--shiki-default-font-weight);text-decoration: var(--shiki-default-text-decoration);}html .dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html.dark .shiki span {color: var(--shiki-dark);background: var(--shiki-dark-bg);font-style: var(--shiki-dark-font-style);font-weight: var(--shiki-dark-font-weight);text-decoration: var(--shiki-dark-text-decoration);}html pre.shiki code .sZZnC, html code.shiki .sZZnC{--shiki-default:#032F62;--shiki-dark:#9ECBFF}html pre.shiki code .sVt8B, html code.shiki .sVt8B{--shiki-default:#24292E;--shiki-dark:#E1E4E8}html pre.shiki code .sj4cs, html code.shiki .sj4cs{--shiki-default:#005CC5;--shiki-dark:#79B8FF}",{"title":931,"searchDepth":1786,"depth":1786,"links":5056},[5057,5058,5059,5063,5066,5070,5073,5077,5080,5083,5089,5090],{"id":291,"depth":1786,"text":292},{"id":463,"depth":1786,"text":464},{"id":907,"depth":1786,"text":908,"children":5060},[5061,5062],{"id":942,"depth":1792,"text":943},{"id":979,"depth":1792,"text":980},{"id":1188,"depth":1786,"text":1189,"children":5064},[5065],{"id":1400,"depth":1792,"text":1401},{"id":1688,"depth":1786,"text":1689,"children":5067},[5068,5069],{"id":1743,"depth":1792,"text":1744},{"id":1861,"depth":1792,"text":1862},{"id":2013,"depth":1786,"text":2014,"children":5071},[5072],{"id":2061,"depth":1792,"text":2062},{"id":2138,"depth":1786,"text":2139,"children":5074},[5075,5076],{"id":2157,"depth":1792,"text":2158},{"id":2305,"depth":1792,"text":2306},{"id":3065,"depth":1786,"text":3066,"children":5078},[5079],{"id":3203,"depth":1792,"text":3204},{"id":3407,"depth":1786,"text":3408,"children":5081},[5082],{"id":3604,"depth":1792,"text":3605},{"id":4359,"depth":1786,"text":4360,"children":5084},[5085,5086,5087,5088],{"id":4366,"depth":1792,"text":4367},{"id":4719,"depth":1792,"text":4720},{"id":4768,"depth":1792,"text":4769},{"id":4779,"depth":1792,"text":4780},{"id":4952,"depth":1786,"text":4953},{"id":5000,"depth":1786,"text":5001},"\u002Fimages\u002Fcovers\u002Fbeating-llamacpp-from-scratch-consumer-amd.png","2026-10-08","Software Architect Mihai Farcas details engineering Strata: an ultra-fast local LLM inference engine for Qwen 3.5 on AMD Radeon RX 7900 XTX (gfx1100). How custom HIP kernels, split-KV reduction, and zero-allocation Rust beat llama.cpp and Ollama at 83.0 tok\u002Fs BF16 decode.","md",{},"\u002Fblog\u002Fbeating-llamacpp-from-scratch-consumer-amd",{"title":5,"description":5093},"blog\u002Fbeating-llamacpp-from-scratch-consumer-amd",[5100,5101,1774,5102,5103,5104,5105,5106,5107,5108,5109,5110],"local-llm","local-ai","rocm","amd","gpu-kernels","systems-engineering","cuda","qwen","strata","mihai-farcas","ai-infrastructure","blog_post","3iWqQkqLjo-nqjzL893twtin5GIIR7-eBqvOQpKJPeM",null,1791464645839]