Back to blogs

About the author

xbill

@xbill

Master of MCP and herding cats

New York

From the author

Articles by xbill

Explore all insights

Explore xbill's latest articles and ideas.

Aug 16

Shipping a vision-model verdict on Bedrock and Lightsail

One FastAPI route, one Converse call with forced tool use, one container service — and the three AWS details that cost me the most time. 20/20 on the fixture set, median 880 ms against the live deployment.

9 min readRead article

Aug 15

I built a security scanner that checks if you are a dog

A live-video scanner that barks when it sees a dog, built by a self-paced AI loop over one weekend. Every green checkmark in this project was, at some point, green over something broken.

9 min readRead article

Aug 14

Serving Gemma4 with Rust on vLLM 🦀

Step-by-step: getting vLLM's Rust frontend built and running on an aarch64 EC2 G5g box. rustup, setuptools-rust, protoc, the release flag, and how to prove the Rust frontend is actually the one answering.

10 min readRead article

Aug 14

Installing Rust for vLLM on Graviton: a G5g walk-through 🦀

Step-by-step: getting vLLM's Rust frontend built and running on an aarch64 EC2 G5g box. rustup, setuptools-rust, protoc, the release flag, and how to prove the Rust frontend is actually the one answering.

10 min readRead article

Aug 15

Building and Serving vLLM with Rust

This tutorial walks through deploying vLLM and some of the key Rust tools used for building and...

9 min readRead article

Aug 14

Looker's Native MCP Server with Claude Code

Looker hosts its own MCP server now. This walks through connecting Claude Code to it, pairing it with...

9 min readRead article

Aug 13

Running Gemma 4 on EC2 G5g: Graviton2 AMD with NVIDIA GPU

A field report on serving Gemma 4 E2B under vLLM on AWS G5g — the only aarch64 + SM 7.5 hardware there is. No published build covers that combination, AWS quietly solves half of it, and the thing that actually blocks you is 64 KiB of shared memory.

9 min readRead article

Aug 13

Running Gemma 4 on EC2 G5g: Graviton2 AMD with NVIDIA GPU

A field report on serving Gemma 4 E2B under vLLM on AWS G5g — the only aarch64 + SM 7.5 hardware there is. No published build covers that combination, AWS quietly solves half of it, and the thing that actually blocks you is 64 KiB of shared memory.

9 min readRead article

Aug 13

Do I Still Need a Monkey Patch for Gemini Live?

No. And deleting 187 lines of it was the single biggest benefit of moving to ADK 2.x — but it was not...

11 min readRead article

Aug 12

MCP Configuration for Looker with Codex

This article covers the MCP setup and configuration for using Looker with Codex to enhance and extend...

24 min readRead article

Aug 11

The unofficial TPU migration guide: Cloud TPU API to Compute Engine

Cloud TPU resources in Compute Engine puts it plainly: The Cloud TPU API is no longer under...

17 min readRead article

Aug 11

Serving Gemma 4 E2B on a TPU v6e-1: what Trillium buys, and what it doesn't

Serving Gemma 4 E2B on a TPU v6e-1 A Cloud TPU v6e-1 (Trillium) costs 2.25× a v5e-1 and...

20 min readRead article