Back to blogs

About the author

xbill

@xbill

Master of MCP and herding cats

New York

From the author

Articles by xbill

Explore all insights

Explore xbill's latest articles and ideas.

Oct 2

Write Markdown Once, Publish It Everywhere: dev.to, Medium, AWS Builder Center and LinkedIn

Markdown is easy to write and hard to publish. Every destination renders it differently, two have no API, and the failures show up only after you hit Publish. A step by step walk-through of publishing-kit, an agent skill that builds, checks and posts each version, used here to publish this article.

10 min readRead article

Oct 2

Write Markdown Once, Publish It Everywhere: dev.to, Medium, AWS Builder Center and LinkedIn

Markdown is easy to write and hard to publish. Every destination renders it differently, two have no API, and the failures show up only after you hit Publish. A step by step walk-through of publishing-kit, an agent skill that builds, checks and posts each version, used here to publish this article.

10 min readRead article

Oct 2

Repacked QAT Gemma 4 on One TPU v5e: 12B Serves at 675 Tokens per Second

Google's quantization-aware-trained Gemma 4 weights, repacked into int4 and int8 formats vLLM serves on TPU. On one v5e chip the repacks serve every size from E2B to 26B, read the suite level with bf16 through 12B, score up to 2.4 points above Google's own 4-bit exports at the same speed, and put 12B on the chip at 11.31 GiB and 675 output tokens per second.

9 min readRead article

Oct 2

Gemma 4 QAT on One TPU v5e: What Runs and What Doesn't

Google's quantization-aware-trained Gemma 4 weights, repacked into int4 and int8 formats vLLM serves on TPU, on one v5e chip. What runs: every size from E2B to 26B, level with bf16 through 12B, up to 2.4 points above Google's own 4-bit exports at the same speed, and 12B at 675 output tokens per second. What doesn't: bf16 above E2B, 31B in any build, and 26B past a 2,176-token context.

9 min readRead article

Oct 2

Gemma 4 at Over 70 Tokens/s on a 2021 Laptop's 4 GB GPU: The Live Demo, Step by Step

The live demo, step by step: Gemma 4 E2B answering at 76 tok/s from a GTX 1650 Ti with 4 GB of memory, and the exact re-pack of Google's QAT weights on Hugging Face that makes it fit, run faster and stay close to bf16.

12 min readRead article

Sep 30

Gemma 4 on Amazon SageMaker or a VM? The Same Model Server at 1.40x the Price

The same Gemma 4 build, vLLM version and GPU served from a SageMaker endpoint and from a plain EC2 instance, on a T4 and an L4: identical decode and answers, a different call path, and what the managed endpoint's 1.40x buys.

8 min readRead article

Sep 30

Gemma 4 on Amazon SageMaker: The NVIDIA T4 Decodes at 0.8x of the L4 With the Same Answers

Gemma 4's 4-bit builds on SageMaker's smallest GPU, an NVIDIA T4, against the L4: a Turing patch for vLLM, the host image the CUDA 13 container needs, speed, memory, answers and cost per token.

9 min readRead article

Sep 30

Gemma 4 on Amazon SageMaker: The NVIDIA T4 Decodes at 0.8x of the L4 With the Same Answers

Gemma 4's 4-bit builds on SageMaker's smallest GPU, an NVIDIA T4, against the L4: a Turing patch for vLLM, the host image the CUDA 13 container needs, speed, memory, answers and cost per token.

9 min readRead article

Sep 30

Gemma 4 on a Tesla T4, Part 3: Int4 Embeddings Serve E2B in 2.86 GiB at 2.30x bf16

Google's QAT Gemma 4 E2B keeps its embedding tables in bf16, and on a Tesla T4 they are most of the model. Packing them to int4 on the grid QAT trained them onto cuts model loading from 6.33 to 2.86 GiB, with every greedy test output token-identical, and raises vLLM's output throughput 11-37% over Google's own W4A16 export.

11 min readRead article

Sep 30

Gemma 4 on Amazon SageMaker: 4-Bit Embeddings Decode up to 1.39x Faster on One L4

Repacking Gemma 4's QAT weights five ways and serving each on the same SageMaker NVIDIA L4 endpoint: int4 linears, int4 embeddings and lm_head, FP8 and int8, across E2B, E4B, 12B, 26B A4B and 31B.

11 min readRead article

Sep 28

Count It or Compute It: When a Tool Returns Rows, the Models That Count Them Right Spend the Tokens

A Kaggle benchmark of the step agents rarely test: counting what a tool returns. Ten models, 68 questions, one tool that returns the count and one that returns the rows. With the count, every model is right at a flat cost. With the rows, models that reason through the list count 330 ids right and spend 6 to 26 times the tokens doing it; models that answer straight away get 0 to 10 of 21.

10 min readRead article

Sep 26

Google's QAT Gemma 4 26B-A4B on One TPU v6e: 15.6x the KV Cache and 1.9x the Throughput of FP8

Google ships its quantization-aware-trained Gemma 4 26B-A4B as GGUF and as a 48 GiB bf16 export, and as compressed-tensors W4A16 for every size but this one. A lossless repack to W4A16, a W4A16 mixture-of-experts method for vLLM's JAX path on TPU, and one v6e chip: 17.43 GiB of HBM, 53,888 KV tokens and 1,283 output tokens per second, against RedHat's FP8 build at 27.99 GiB, 3,456 tokens and 668. The same checkpoint loads unpatched on vLLM 0.30.0 on an NVIDIA L4.

12 min readRead article