Back to blogs

About the author

xbill

@xbill

Master of MCP and herding cats

New York

From the author

Articles by xbill

Explore all insights

Explore xbill's latest articles and ideas.

Sep 15

Nano Banana 2 Lite in Kiro CLI 3: MCP 2.0, the New Interactions API, and Headless Permissions

The Kiro edition of the Nano Banana 2 Lite MCP server, updated: FastMCP is now MCPServer, google-genai 1.x gets a 400 from the Interactions API, and Kiro CLI 3 needs a permissions rule before it will call the tools headless.

12 min readRead article

Sep 14

Build in the VM, Think on the Mac GPU: Debian 13 on Apple container With a Local Gemma 4

Apple's container CLI 1.4.1 cannot boot the official debian:13 image as a machine, because it has no /sbin/init. A small Dockerfile fixes that and gives you a persistent Debian 13 VM with systemd, your Mac user and your home folder. The VM has no GPU, so Ollama runs on macOS and the VM calls it at 192.168.64.1:8000. From inside Debian, gemma4:e2b answered at 45.2 tok/s, loaded 100% on the M3 GPU.

16 min readRead article

Sep 13

Nano Banana 2 Lite, Revisited: MCP 2.0, the New Interactions API, and Three Agent CLIs

The Nano Banana 2 Lite MCP server from July, updated: FastMCP is now MCPServer, google-genai 1.x gets a 400 from the Interactions API, and one server now runs in Claude Code, Codex and Antigravity CLI.

12 min readRead article

Sep 11

Two Rust Clients for Gemma 4: Calling the Endpoint vs. Calling the MCP Server 🦀

Step by step: a Rust HTTP client (reqwest) and a Rust MCP client (rmcp) ask Gemma 4 E2B the same question on a local llama.cpp GPU and on Cloud Run — what each one can see, and what each one costs.

17 min readRead article

Sep 11

FastMCP Is Now MCPServer on AWS: Moving a boto3 EC2 MCP Server to the MCP Python SDK 2.x

What basic MCP Python SDK 2.x support takes for an AWS EC2 MCP server built on boto3: one rename in server.py, snake_case in the tests, a floor in requirements.txt, and why neither the wire format nor the boto3 calls change.

9 min readRead article

Sep 11

FastMCP Is Now MCPServer: Migrating a Python MCP Server to the MCP SDK 2.x

Step-by-step: moving a FastMCP server to the MCP Python SDK 2.x — what broke, what changed, what did not, and a redeploy of its Gemma 4 vLLM backend to a Cloud Run L4 GPU.

12 min readRead article

Sep 10

Gemma 4 on an 2021 4 GB Laptop GPU: QAT Takes It From 9.5 GiB to 1.6

Step-by-step: running Google's quantization-aware-trained Gemma 4 E2B on a 10th-gen Core i7 laptop with a 4 GB GTX 1650 Ti — why bf16 and int8 cannot fit, why the QAT GGUF does with room to spare, and managing it with an MCP server.

13 min readRead article

Sep 10

2B Gemma 4 Deployment with Cloud Run, NVIDIA L4, MCP SDK 2.x, and Claude Code

Step by step deployment of Gemma 4 E2B to a Cloud Run NVIDIA L4 GPU with vLLM, managed by a Python MCP server migrated to the MCP SDK 2.x.

13 min readRead article

Sep 9

Four Debian 13 Boxes, One Brief: 1,923 Packages on Metal, 328 in the Cloud

The same Debian 13 on a laptop and on AWS, GCE and Azure. The cloud images ship 328-350 packages against 1,923 on metal, no firmware package at all, and no tool to read the NVMe they all boot from. Each machine was then given the same brief and asked to analyse itself. Three of them agreed. The fourth found a partition table its own resize had broken.

11 min readRead article

Sep 9

Four Debian 13 Boxes, One Brief: 1,923 Packages on Metal, 328 in the Cloud

The same Debian 13 on a laptop and on AWS, GCE and Azure. The cloud images ship 328-350 packages against 1,923 on metal, no firmware package at all, and no tool to read the NVMe they all boot from. Each machine was then given the same brief and asked to analyse itself. Three of them agreed. The fourth found a partition table its own resize had broken.

11 min readRead article

Sep 8

The Desktop Looked Right: 2,093 Parse Errors a Boot, and 10.9 Seconds That Weren't Doing Anything

A Debian desktop that looked and behaved correctly was throwing 2,093 CSS parse errors per boot and spending 10.9 seconds of a 39-second boot waiting on nothing. Neither is visible on screen. Both are in the journal, and both were introduced by the script that was supposed to set the machine up.

16 min readRead article

Sep 8

The Cable Buys Headroom: 91% of a USB 2.0 Bus, 3.6% of a Thunderbolt One

The best USB 2.0 pass in 45 tethering records used 91% of what that bus can usably carry. The same link on a SuperSpeed cable uses 3.6%. USB speed is autonegotiated between host, cable and device, the slowest one wins, and nothing anywhere tells you the cable capped you.

18 min readRead article