Back to blogs

About the author

xbill

@xbill

Master of MCP and herding cats

New York

From the author

Articles by xbill

Explore all insights

Explore xbill's latest articles and ideas.

Jul 15

My Demo Script Found a Production Bug on Its First Run: A Tiny Post-Mortem

How a narrated demo for an MCP image server caught an API-contract drift that 10 passing unit tests never could — root cause, fix, and lessons about mocked tests.

5 min readRead article

Jul 15

Build One AI Tool Server, Call It From Three Different Agents (MCP Explained)

A beginner-friendly tour of the Model Context Protocol: one Python server that generates images with Gemini, called by Claude Code, a Google ADK agent, and a Rust CLI.

6 min readRead article

Jul 15

Why did my benchmark stop at N=22? A debugging story in nine bugs

Submission for DEV's Summer Bug Smash — Smash Stories track. There was a file in my repo called...

4 min readRead article

Jul 15

My benchmark's Python column was N/A for a year — CPython's 4300-digit limit, and eight other bugs

Submission for DEV's Summer Bug Smash — Clear the Lineup track. The...

4 min readRead article

Jul 15

TPU Deployments with Gemma 31B, v6e-8, and Antigravity CLI

This article provides a step by step debugging guide for deploying Gemma 4 to a Google Cloud TPU...

27 min readRead article

Jul 15

26B Gemma 4 QAT Deployment with GCE g2-standard, NVIDIA L4, MCP, and Antigravity CLI

This article provides a step by step deployment guide for Gemma 4 to a Google Compute Engine hosted...

17 min readRead article

Jul 14

4B Gemma 4 QAT Deployment with GCE, NVIDIA L4, MCP, and Antigravity CLI

This article provides a step by step deployment guide for Gemma 4 to a Google Compute Engine hosted...

20 min readRead article

Jul 13

Porting Gemma-4 (2B / 4B / 12B) to AWS Inferentia2

A field report on running Google's Gemma-4 on AWS Inferentia2: mixed attention heads, the vLLM / optimum-neuron / NxD dead-ends, and the neuronx-cc compiler limits.

16 min readRead article

Jul 13

Porting Gemma-4 (2B / 4B / 12B) to AWS Inferentia2

A field report on running Google's Gemma-4 on AWS Inferentia2: mixed attention heads, the vLLM / optimum-neuron / NxD dead-ends, and the neuronx-cc compiler limits.

16 min readRead article

Jul 13

2B Gemma 4 QAT Deployment with GCE, NVIDIA L4, MCP, and Antigravity CLI

This article provides a step by step deployment guide for Gemma 4 to a Google Compute Engine hosted...

15 min readRead article

Jul 12

TPU Deployments with Gemma 31B, v6e-4, and Antigravity CLI

This article provides a step by step debugging guide for deploying Gemma 4 to a Google Cloud TPU...

26 min readRead article

Jul 11

Leetcode for the win!

GRIND404: I turned my "Passion" for LeetCode into a playable arcade game DEV Weekend...

1 min readRead article