About the author
xbill
@xbill
Master of MCP and herding cats
New York
Oct 2
Write Markdown Once, Publish It Everywhere: dev.to, Medium, AWS Builder Center and LinkedIn
Markdown is easy to write and hard to publish. Every destination renders it differently, two have no API, and the failures show up only after you hit Publish. A step by step walk-through of publishing-kit, an agent skill that builds, checks and posts each version, used here to publish this article.
Oct 2
Write Markdown Once, Publish It Everywhere: dev.to, Medium, AWS Builder Center and LinkedIn
Markdown is easy to write and hard to publish. Every destination renders it differently, two have no API, and the failures show up only after you hit Publish. A step by step walk-through of publishing-kit, an agent skill that builds, checks and posts each version, used here to publish this article.
Oct 2
Repacked QAT Gemma 4 on One TPU v5e: 12B Serves at 675 Tokens per Second
Google's quantization-aware-trained Gemma 4 weights, repacked into int4 and int8 formats vLLM serves on TPU. On one v5e chip the repacks serve every size from E2B to 26B, read the suite level with bf16 through 12B, score up to 2.4 points above Google's own 4-bit exports at the same speed, and put 12B on the chip at 11.31 GiB and 675 output tokens per second.
Oct 2
Gemma 4 QAT on One TPU v5e: What Runs and What Doesn't
Google's quantization-aware-trained Gemma 4 weights, repacked into int4 and int8 formats vLLM serves on TPU, on one v5e chip. What runs: every size from E2B to 26B, level with bf16 through 12B, up to 2.4 points above Google's own 4-bit exports at the same speed, and 12B at 675 output tokens per second. What doesn't: bf16 above E2B, 31B in any build, and 26B past a 2,176-token context.
Oct 2
Gemma 4 at Over 70 Tokens/s on a 2021 Laptop's 4 GB GPU: The Live Demo, Step by Step
The live demo, step by step: Gemma 4 E2B answering at 76 tok/s from a GTX 1650 Ti with 4 GB of memory, and the exact re-pack of Google's QAT weights on Hugging Face that makes it fit, run faster and stay close to bf16.
Sep 30
Gemma 4 on Amazon SageMaker or a VM? The Same Model Server at 1.40x the Price
The same Gemma 4 build, vLLM version and GPU served from a SageMaker endpoint and from a plain EC2 instance, on a T4 and an L4: identical decode and answers, a different call path, and what the managed endpoint's 1.40x buys.
Sep 30
Gemma 4 on Amazon SageMaker: The NVIDIA T4 Decodes at 0.8x of the L4 With the Same Answers
Gemma 4's 4-bit builds on SageMaker's smallest GPU, an NVIDIA T4, against the L4: a Turing patch for vLLM, the host image the CUDA 13 container needs, speed, memory, answers and cost per token.
Sep 30
Gemma 4 on Amazon SageMaker: The NVIDIA T4 Decodes at 0.8x of the L4 With the Same Answers
Gemma 4's 4-bit builds on SageMaker's smallest GPU, an NVIDIA T4, against the L4: a Turing patch for vLLM, the host image the CUDA 13 container needs, speed, memory, answers and cost per token.
Sep 30
Gemma 4 on a Tesla T4, Part 3: Int4 Embeddings Serve E2B in 2.86 GiB at 2.30x bf16
Google's QAT Gemma 4 E2B keeps its embedding tables in bf16, and on a Tesla T4 they are most of the model. Packing them to int4 on the grid QAT trained them onto cuts model loading from 6.33 to 2.86 GiB, with every greedy test output token-identical, and raises vLLM's output throughput 11-37% over Google's own W4A16 export.
Sep 30
Gemma 4 on Amazon SageMaker: 4-Bit Embeddings Decode up to 1.39x Faster on One L4
Repacking Gemma 4's QAT weights five ways and serving each on the same SageMaker NVIDIA L4 endpoint: int4 linears, int4 embeddings and lm_head, FP8 and int8, across E2B, E4B, 12B, 26B A4B and 31B.
Sep 28
Count It or Compute It: When a Tool Returns Rows, the Models That Count Them Right Spend the Tokens
A Kaggle benchmark of the step agents rarely test: counting what a tool returns. Ten models, 68 questions, one tool that returns the count and one that returns the rows. With the count, every model is right at a flat cost. With the rows, models that reason through the list count 330 ids right and spend 6 to 26 times the tokens doing it; models that answer straight away get 0 to 10 of 21.
Sep 26
Google's QAT Gemma 4 26B-A4B on One TPU v6e: 15.6x the KV Cache and 1.9x the Throughput of FP8
Google ships its quantization-aware-trained Gemma 4 26B-A4B as GGUF and as a 48 GiB bf16 export, and as compressed-tensors W4A16 for every size but this one. A lossless repack to W4A16, a W4A16 mixture-of-experts method for vLLM's JAX path on TPU, and one v6e chip: 17.43 GiB of HBM, 53,888 KV tokens and 1,283 output tokens per second, against RedHat's FP8 build at 27.99 GiB, 3,456 tokens and 668. The same checkpoint loads unpatched on vLLM 0.30.0 on an NVIDIA L4.