Purify field notes / Vol. 01
The web, after the noise.
Working notes on extraction, search, agents, and the operational details between a URL and useful data.
Index / 06 notes
Read from the source.
Web scraping for RAG: a complete guide
Build a web-data pipeline for retrieval with clean content, evidence-preserving chunks, refresh policy, and evaluation at every stage.
Why I rewrote Firecrawl in Go
The architectural choices behind Purify: a Go service, an HTTP-first path, browser fallback, and one system for scrape, crawl, map, and extract.
Best web scraping API for AI agents in 2026
A practical guide to evaluating web-data APIs for agent workflows, including integration, response contracts, latency visibility, and failure handling.
Best Firecrawl alternatives in 2026
A practical framework for comparing Firecrawl, Crawl4AI, Jina Reader, and Purify without relying on a stale price sheet or an unreproducible leaderboard.
How to reduce AI token costs when scraping the web
A practical method for measuring how web-content cleaning changes model input, without treating one recorded run as a universal benchmark.
How to set up an MCP server for web scraping
Connect an MCP-compatible client to the Purify MCP binary using the same configuration contract shown in the dashboard.