About the author
xbill
@xbill
Master of MCP and herding cats
New York
Jul 23
Teaching Kiro to Paint: A Stateful Image-Editing Skill Built on Gemini's Interactions API and MCP
TL;DR: nb2lite-skill-kiro wraps Google's gemini-3.1-flash-lite-image model (NB2Lite) in a tiny...
Jul 22
Teaching Claude Code to Paint: A Stateful Image-Editing Skill Built on Gemini's Interactions API and MCP
How nb2lite-skill-claude packages Google's gemini-3.1-flash-lite-image as a Claude Code skill + MCP server — with multi-turn stateful edits, an idiot-proof install guide, and a dogfooded cover image.
Jul 21
Gemma 4 E2B on a Single TPU v6e Chip: A Serving Deep Dive
What it took to deploy, why the QAT checkpoints refuse to load, and what one flex-start v6e chip is actually worth — measured live.
Jul 21
tpu-management: a Claude Code skill for running Gemma 4 on Cloud TPUs
A Claude Code skill + MCP server that provisions Google Cloud TPU capacity, serves Gemma 4 with vLLM, benchmarks it, and tears it down — installable in one command.
Jul 20
Gemma4 DevOps In Action
In the last entry I got Gemma-4's 128-expert MoE running on an inf2.24xlarge and signed off with...
Jul 20
Inferntia2 DevOps in Action
In the last entry I got Gemma-4's 128-expert MoE running on an inf2.24xlarge and signed off with...
Jul 17
Five Gemma-4 models, one accelerator: what porting E2B 31B to AWS Inferentia2 taught me
I ported the whole Gemma-4 family — E2B, E4B, 12B, 31B, and the 26B-A4B MoE — to run on...
Jul 17
Porting Gemma-4 12B (the encoder-free multimodal one) to AWS Inferentia2
The 12B ships as a multimodal class with no encoder loaded, and its sliding-window attention overflows Neuron's fused-attention SBUF. Three surgical fixes and it serves Paris at ~15 tok/s on one inf2.8xlarge.
Jul 17
My Inferentia port matched its reference token-for-token — and still output garbage
Porting Gemma-4 31B (dense) to AWS Inferentia2: the tensor-parallel recipe that worked at 12B collapses at 31B, NxD ModelBuilder saves it — and then a passing validation lies to your face.
Jul 17
Porting a 128-expert MoE (Gemma-4 26B-A4B) to AWS Inferentia2 — where every rank weighted the wrong experts
The MoE was the hard one: a dual-path FFN, a sparse expert loop that won't trace, and a bug where the device output was empty while the CPU reference was perfect and every unit test passed.
Jul 17
Five Gemma-4 models, one accelerator: what porting E2B 31B to AWS Inferentia2 taught me
A side-by-side of all five Gemma-4 variants on Inferentia2 — PLE, KV-sharing, MatFormer, mixed attention, and a 128-expert MoE — the recipe that evolved to carry all of them, and the one bug that shows up in every single one.
Jul 15
Smash Story: The Demo Script That Out-Debugged My Test Suite
A green 10-test suite, a broken production default, and the 10-minute smash — how a live demo caught an API-contract bug that mocks never could.