Back to blogs

About the author

xbill

@xbill

Master of MCP and herding cats

New York

From the author

Articles by xbill

Explore all insights

Explore xbill's latest articles and ideas.

Jul 23

Teaching Kiro to Paint: A Stateful Image-Editing Skill Built on Gemini's Interactions API and MCP

TL;DR: nb2lite-skill-kiro wraps Google's gemini-3.1-flash-lite-image model (NB2Lite) in a tiny...

8 min readRead article

Jul 22

Teaching Claude Code to Paint: A Stateful Image-Editing Skill Built on Gemini's Interactions API and MCP

How nb2lite-skill-claude packages Google's gemini-3.1-flash-lite-image as a Claude Code skill + MCP server — with multi-turn stateful edits, an idiot-proof install guide, and a dogfooded cover image.

8 min readRead article

Jul 21

Gemma 4 E2B on a Single TPU v6e Chip: A Serving Deep Dive

What it took to deploy, why the QAT checkpoints refuse to load, and what one flex-start v6e chip is actually worth — measured live.

8 min readRead article

Jul 21

tpu-management: a Claude Code skill for running Gemma 4 on Cloud TPUs

A Claude Code skill + MCP server that provisions Google Cloud TPU capacity, serves Gemma 4 with vLLM, benchmarks it, and tears it down — installable in one command.

3 min readRead article

Jul 20

Gemma4 DevOps In Action

In the last entry I got Gemma-4's 128-expert MoE running on an inf2.24xlarge and signed off with...

7 min readRead article

Jul 20

Inferntia2 DevOps in Action

In the last entry I got Gemma-4's 128-expert MoE running on an inf2.24xlarge and signed off with...

7 min readRead article

Jul 17

Five Gemma-4 models, one accelerator: what porting E2B 31B to AWS Inferentia2 taught me

I ported the whole Gemma-4 family — E2B, E4B, 12B, 31B, and the 26B-A4B MoE — to run on...

6 min readRead article

Jul 17

Porting Gemma-4 12B (the encoder-free multimodal one) to AWS Inferentia2

The 12B ships as a multimodal class with no encoder loaded, and its sliding-window attention overflows Neuron's fused-attention SBUF. Three surgical fixes and it serves Paris at ~15 tok/s on one inf2.8xlarge.

4 min readRead article

Jul 17

My Inferentia port matched its reference token-for-token — and still output garbage

Porting Gemma-4 31B (dense) to AWS Inferentia2: the tensor-parallel recipe that worked at 12B collapses at 31B, NxD ModelBuilder saves it — and then a passing validation lies to your face.

8 min readRead article

Jul 17

Porting a 128-expert MoE (Gemma-4 26B-A4B) to AWS Inferentia2 — where every rank weighted the wrong experts

The MoE was the hard one: a dual-path FFN, a sparse expert loop that won't trace, and a bug where the device output was empty while the CPU reference was perfect and every unit test passed.

6 min readRead article

Jul 17

Five Gemma-4 models, one accelerator: what porting E2B 31B to AWS Inferentia2 taught me

A side-by-side of all five Gemma-4 variants on Inferentia2 — PLE, KV-sharing, MatFormer, mixed attention, and a 128-expert MoE — the recipe that evolved to carry all of them, and the one bug that shows up in every single one.

6 min readRead article

Jul 15

Smash Story: The Demo Script That Out-Debugged My Test Suite

A green 10-test suite, a broken production default, and the 10-minute smash — how a live demo caught an API-contract bug that mocks never could.

4 min readRead article