Case study · Founder / Developer

GAScout

Built because Georgia's delinquent-tax data exists only as a 9,915-page fixed-width mainframe PDF — a format built for a printer, not for analysis. The extraction engine's job is to turn that dump into something a policy researcher can query, cut, and trust: parsed and checksum-verified line by line, with zero dropped rows, so aggregate figures are traceable back to source instead of estimated.

2026livePython · Astro v5 · Vanilla JavaScript · Cloudflare Pages · pypdf
$62.75Muncollected delinquent tax liability tracked across 8 tax cycles
409,142line items parsed from 9,915 mainframe PDF pages with 0 dropped rows
90automated offline test fixtures verifying parser & checksum integrity

A static public-records intelligence dashboard and daily extraction pipeline built on DeKalb County delinquent tax listings.

County tax listings are published as a monolithic 9,915-page fixed-width mainframe PDF file (DQ205GADEK). The python extraction engine ingests the document offline line-by-line, detecting and auto-correcting 32 column-overflow shifts via a 5-column checksum verification algorithm. The pipeline enforces 100% line accounting — zero dropped rows, zero malformed records — backed by 90 automated offline test fixtures.

The web interface presents aggregate financial insights without publishing personal PII: