Back to blogs

About the author

xbill

@xbill

Master of MCP and herding cats

New York

From the author

Articles by xbill

Explore all insights

Explore xbill's latest articles and ideas.

Sep 5

Generosity Is a Default Setting

A USB tethering diagnostic that costs no cellular data to run, forty measurements across fourteen phones, and one free setting that took the worst run from 15 to 106 Mbps.

9 min readRead article

Sep 4

AWS Has Two Iceberg REST Catalogs: What Each One Actually Serves

Glue and S3 Tables both implement the Apache Iceberg REST catalog specification. One identical request suite against both shows the same totals and thirteen behavioural differences, including two that require opposite things of the same drop request.

8 min readRead article

Sep 4

Seven Iceberg REST Catalogs: What They Declare, and What They Serve

One request suite against seven Apache Iceberg REST catalog implementations — Polaris, BigLake, Glue, S3 Tables, Unity, Horizon and OneLake — comparing what each catalog declares it supports against what it actually serves.

17 min readRead article

Sep 3

ChromeOS Lookalikes, Two Ways: One With Drivers, One Without

chromeos-boot holds two unrelated scripts under one name: stage seeds a real Crostini container from a private bucket, flex skins a bare-metal Debian desktop to look like one. The split exists because Crostini's guest kernel can't load the NVIDIA driver.

8 min readRead article

Sep 1

Three Gemma 4 Deployments on One T4G for Under $3: What the Runtime Changes, and What It Doesn't

vLLM, JAX and PyTorch serving the same Gemma 4 E2B checkpoint on the same AWS G5g GPU, on one harness and one statistic. The decode ranking reverses on boot time. Nineteen instances, four and a half instance-hours, under $3 - which is what let five wrong claims get caught.

14 min readRead article

Sep 1

Three Gemma 4 Deployments on One T4G for Under $3: What the Runtime Changes, and What It Doesn't

vLLM, JAX and PyTorch serving the same Gemma 4 E2B checkpoint on the same AWS G5g GPU, on one harness and one statistic. The decode ranking reverses on boot time. Nineteen instances, four and a half instance-hours, under $3 - which is what let five wrong claims get caught.

14 min readRead article

Sep 1

Streamline Publishing with a Claude Code Skill

A Claude Code skill that turns one markdown file into dev.to, AWS Builder Center, Medium and LinkedIn versions, checks them before they ship, and posts the ones with an API — plus the debugging tools for when a destination mangles something.

7 min readRead article

Sep 1

Streamline Publishing with a Claude Code Skill

A Claude Code skill that turns one markdown file into dev.to, AWS Builder Center, Medium and LinkedIn versions, checks them before they ship, and posts the ones with an API — plus the debugging tools for when a destination mangles something.

7 min readRead article

Aug 31

g5g vs g6 for LLM Serving: the Same Code, and 3.7x the Throughput

Serving Gemma 4 E2B in pure JAX on AWS g5g.2xlarge and g6.2xlarge with a byte-identical payload. The older instance loses 87% of decode to dtype conversion, and nothing in the logs says so.

6 min readRead article

Aug 31

Gemma 4 in Pure JAX: What Changes Between Turing and Ada, and What Doesn't

One hand-written Gemma 4 port, no PyTorch and no vLLM, on two NVIDIA GPUs a generation apart. Most of it ports untouched. Two things do not, and one of them was quietly eating 87% of decode.

8 min readRead article

Aug 31

The Cheapest CUDA GPU on AWS Has an Arm CPU — and You Probably Want the Intel One

Gemma 4 E2B on the two cheapest whole-GPU CUDA instances AWS sells. Same Turing GPU, different host CPU. One is cheaper per hour, the other is cheaper per token, and the reason is not the CPU.

11 min readRead article

Aug 29

Gemma 4 in Pure JAX: What Ports from TPU to GPU, and What Doesn't

This article is about running a hand-written Gemma 4 port in pure JAX on three different...

6 min readRead article