POST-based paginated download with offset or OID-range strategy auto-detected per layer, a thread pool across layers, retries, and atomic temp-then-rename writes so a failure never leaves a half-file behind.
Onboard a city by writing a config file.
Municipal GIS is a thousand different schemas wearing the same shape. This pipeline reads a jurisdiction's endpoints out of a YAML file, downloads them, translates every column into a fixed canonical vocabulary, and runs the same spatial analysis it runs everywhere else — producing the opportunity layers that Land Intelligence and the Execution Layer consume.
Every city is a rewrite
The analysis a developer wants is the same in every market: what is underbuilt, what is vacant, what is distressed, what can be assembled. The data underneath it never is.
Columns never match
Miami calls it FOLIO. Austin calls it PROP_ID. NYC calls it a BBL. Hard-code one and the query is worth nothing in the next county.
Coverage is uneven
One city publishes liens and recertification records; the next publishes neither. A pipeline that requires a complete dataset never runs anywhere.
Extraction is slow and fragile
Paginated REST endpoints time out, cap result counts, and disagree about whether they support offsets at all. Half-written files corrupt the next stage.
The stack gets heavy
The obvious answer is a geopandas / GDAL / PostGIS deployment for something that should be a scheduled job producing a folder of files.
Five stages, driven entirely by the config file
The YAML names the endpoints, the field mappings and the outputs. Nothing else in the system changes when a jurisdiction is added.
DuckDB sorts features along an ST_Hilbert space-filling curve and writes GeoParquet 1.1 with a bbox struct and ZSTD compression, so spatially local reads touch a handful of row groups.
One DuckDB view per semantic role, generated from the config's field map. The parquet keeps its native column names; the view does the translation at query time.
Fifteen city-agnostic spatial SQL queries run against the canonical views. Each writes a GeoJSON layer carrying source URLs and attribution on every feature.
A manifest describing each produced layer — title, geometry, styling, legend metric, exposed properties and upstream attribution — which the map application reads to build its UI.
One vocabulary, translated per city
Every analytical query in the system references canonical column names only —
parcel_id, lot_size,
zone_code, far. A jurisdiction's config
maps those names onto whatever the county actually called them, and the pipeline emits a
view that does the aliasing and the type coercion in one place.
- Native data is never rewritten. The parquet keeps the county's original column names, so provenance survives all the way to the output.
- Views are virtual — no duplication, no second copy of a 595,000-row parcel table.
- Type coercion lives in the view, not scattered through the queries, and uses a non-throwing cast so a stray null in a source column cannot take a query down.
- Adding a city adds a mapping. The analytical SQL is not touched, and a test asserts that no query can smuggle a city-specific column name back in.
NULL for it rather than failing. Queries that depend on
it degrade; every other query still runs. Partial configs are valid configs — which is what
makes onboarding a city with thin open data possible at all.
Thirteen roles a jurisdiction can fill
A config binds each role to one or more real layers. No city has all thirteen — and none needs to. Query availability is derived from which roles a jurisdiction actually bound.
| Role | Geometry | What it feeds |
|---|---|---|
parcels | Polygon | Underbuilt, vacant, rezoning, assemblage, transit proximity, composite scoring |
zoning | Polygon | Underbuilt, rezoning, adaptive reuse, high-density near transit |
zoning_overlays | Polygon | Adaptive reuse |
future_land_use | Polygon | Rezoning candidates |
building_footprints | Polygon | Underbuilt (FAR utilisation), tallest buildings |
permits | Point | Stalled permits, active new construction, composite scoring |
violations | Point | Open violations, composite scoring |
liens | Point | Lien hotspots, composite scoring |
recertification | Point | Recertification leads and clusters, composite scoring |
transit_stations | Point | Transit proximity, high-density zoning near transit |
transit_lines | Polyline | Reference context |
water_service | Polygon | Vacant developable land |
sewer_service | Polygon | Vacant developable land |
* A single role can merge several upstream layers into one view — Miami's transit role, for example, unions the Metrorail and Metromover station services and tags each with its system name.
Fifteen analyses, written once
Each produces a GeoJSON layer with its own property set, source URLs and upstream attribution. These are the opportunity layers that appear in Land Intelligence and the Execution Layer's marketplace.
| Layer | Roles it needs | What it finds |
|---|---|---|
| Underbuilt parcels | parcels, zoning, footprints | Large lots in non-low-rise zones built out to well under the floor-area ratio their zoning already permits |
| Vacant developable land | parcels, water, sewer | Vacant parcels that already sit inside both a water and a sewer service area — the cheapest ground-up sites |
| Rezoning candidates | parcels, zoning, future land use | Parcels whose future land use designation permits far more density than the current zoning does |
| Assemblage candidates | parcels | Adjacent parcels under different ownership that combine into a materially larger development footprint |
| Adaptive reuse candidates | parcels, zoning | Buildings whose current use no longer matches what the district is zoned for |
| Transit proximity parcels | parcels, transit stations | Parcels inside a configurable walk-distance buffer of rapid transit — the TOD opportunity zone |
| High-density zoning near transit | zoning, transit stations | Districts that already allow the tallest, densest development and sit on transit |
| Recertification leads | recertification | Buildings approaching or past their 40/50-year recertification, tiered by urgency against the current year |
| Recertification clusters | recertification | Multi-building addresses facing several recertification deadlines at once |
| Open violations | violations | Active code violations, ranked by escalation tier — notice, lien, then lien plus overdue recertification |
| Lien hotspots | liens, parcels | Parcels accumulating multiple compliance liens — an off-market distress signal |
| Stalled permits | permits | Permits still open long after issue, with the issue date exposed so the map can calculate staleness live |
| Active new construction | permits | Live new-build, demolition and major-alteration permits — the current development pipeline |
| Tallest buildings | building footprints | A reference benchmark set by reported height |
| Composite opportunity score | parcels adaptive | A weighted multi-signal score over violations, liens, pending recertification and active permits |
* The composite score adapts its own SQL to the roles a jurisdiction bound — a city without a liens layer still gets a score, computed from the signals it does publish. Owner and address exclusion filters are configurable per query, so public conservation land, parks, rights of way and utility plots can be kept out of the results.
Find the endpoints, map the fields, run it
No code change, no deployment, no schema migration — a new file and a command.
schema_version: 1 jurisdiction: name: "City of Austin" state: TX county: "Travis" centroid: { lat: 30.2672, lon: -97.7431 } sources: parcels: url: https://services.arcgis.com/…/FeatureServer/0 layer_name: austin_parcels srid: 2277 # reprojected to 4326 on the way in field_map: parcel_id: PROP_ID address: SITE_ADDR owner: OWNER_NAME lot_size: SQ_FT zoning: url: https://services.arcgis.com/…/FeatureServer/5 layer_name: austin_zoning field_map: zone_code: ZONING_TYPE far: FAR_MAX max_height: HEIGHT_MAX queries: # optional tuning per city underbuilt_parcels: { min_lot_sqft: 10000 } transit_proximity_parcels: { buffer_meters: 800 }
# the whole pipeline python -m opportunity_pipeline.run \ jurisdictions/austin_tx.yaml # or one stage at a time python -m opportunity_pipeline.run \ jurisdictions/austin_tx.yaml --stages extract,convert # re-run only the analysis against parquet already on disk python -m opportunity_pipeline.run \ jurisdictions/austin_tx.yaml --stages analyze # what lands on disk data/austin_tx/ ├── geojson/ # raw extraction, native columns ├── geoparquet/ # Hilbert-sorted, ZSTD, bbox struct ├── opportunities/ # the analysis layers │ ├── 15_underbuilt_parcels.geojson │ ├── 16_assemblage_candidates.geojson │ └── 17_composite_opportunity.geojson └── manifest.json # the map application's contract
- Queries skip, they do not crash. Every analysis declares the roles it requires; the engine checks what a jurisdiction bound and quietly skips the rest.
- Analysis is resumable. Existing outputs are left alone unless forced, so regenerating one layer means deleting one file and re-running the stage.
- Extraction is re-entrant. Files are written to a temporary name and renamed on completion, so an interrupted run never poisons the next one.
- Configs are validated before anything downloads — endpoints, field mappings and required columns are checked, and the result is a report you can read rather than a stack trace.
{
"jurisdiction": "miami_fl",
"generated_at": "2026-08-04T06:00:00Z",
"layers": [
{
"id": "underbuilt_parcels",
"title": "Underbuilt Parcels",
"type": "geojson",
"url": "opportunities/15_underbuilt_parcels.geojson",
"geometry": "polygon",
"style": { "fill-color": "#e8a33d" },
"legend": {
"metric": "far_ratio",
"label": "Built FAR / Allowed FAR"
},
"properties": ["parcel_id", "address",
"zone_code", "far_ratio"]
}
],
"attribution": ["Miami-Dade County Property Appraiser",
"City of Miami — Miami 21"]
}
Three markets, three completely different data regimes
The point of the abstraction is that these three share no column names, no service conventions and — in New York's case — not even the same kind of endpoint.
Miami, FL
The reference jurisdiction, and the deepest coverage — parcels, Miami 21 zoning, future land use, building footprints, permits, code violations, liens, 40-year recertification, Metrorail and Metromover, and water and sewer service areas.
- All fifteen analyses available
- Miami-Dade County + City of Miami GIS
Fort Lauderdale, FL
A second Florida city on an entirely separate GIS server with its own naming conventions — tax parcels, zoning districts, proposed land use, building footprints, permits and service requests standing in for the violations role.
- Onboarded by config alone
- Broward County + City of Fort Lauderdale GIS
New York, NY
The hard case. MapPLUTO and zoning districts come from ArcGIS; DOB permits, DOB and HPD violations and OATH ECB liens come from the city's open-data API instead — different transport, same canonical vocabulary downstream.
- Mixed ESRI and open-data sources
- NYC DCP, DOB, HPD, DoITT, MTA
* Coverage per city is bounded by what that jurisdiction publishes — Fort Lauderdale, for example, produces no lien or recertification layers because no such service exists to bind. Ask which analyses are available for the markets you care about.
Half of zoning is prose, so it gets its own pipeline
Setbacks, permitted uses, dimensional standards and parking minimums live in ordinance text, not in a feature service. A companion pipeline discovers a city's code portal, pulls the zoning chapters, and stores them in a form the map and the AI Analyst can actually query — which is what backs ordinance search in Land Intelligence.
- Portal discovery — Municode, American Legal Publishing, eCode360 and Qcode are probed by URL pattern first, at no cost, with web search and a language model only as fallback.
- Chapter navigation — the table of contents is read to find the zoning chapters, with a keyword-only mode that skips model calls entirely.
- Clean Markdown output per chapter, plus optional structured extraction into zone districts, permitted uses, dimensional standards, parking requirements and overlay districts.
- Search-ready storage — a full-text index with stemming and BM25 ranking, plus overlapping content chunks sized for retrieval, so an answer can be cited back to the section it came from.
It behaves on other people's servers
Municipal infrastructure is not built for scraping, and a pipeline that gets an organisation blocked is worse than no pipeline.
- Concurrency is capped globally, with a deliberate delay after every successful fetch.
- Batch runs are sequential across cities by design — parallelising them is a rate-limit posture decision, not an optimisation.
- Results are cached, so a re-run costs nothing until you explicitly force a refresh.
- Pagination is followed with a hard cap, so a badly behaved portal cannot turn one city into an unbounded crawl.
* Four runtime dependencies: a spatial SQL engine, an HTTP client, a config validator and a YAML parser. No geopandas, no fiona, no shapely, no GDAL install and no spatial database to operate — deliberately, so the pipeline runs as a scheduled job on ordinary infrastructure.
Shipped, and what is next
Stated plainly, because the difference matters when you are deciding what to build on.
| Capability | Status | What it means |
|---|---|---|
| Config-driven pipeline, canonical layer, 15 analyses | Shipped | Running against live municipal endpoints with a regression suite over the canonical SQL and the reference jurisdiction |
| Onboarding a city by hand | Shipped | Proven twice beyond the reference city, including one on a completely different class of endpoint |
| Assisted discovery Planned | In design | Crawl a city's service directory, classify each layer into a semantic role and draft the field mapping for a human to review — the crawl deterministic, the judgement calls assisted |
| Scheduled orchestration Planned | In design | Nightly refreshes per jurisdiction with caching keyed on upstream change, retries and run observability |
| Vector tile output Planned | In design | Automatic switch to single-file vector tiles when a layer outgrows what GeoJSON should carry, with the manifest declaring which form a layer took |
| Text and geospatial tracks joined Planned | In design | Each zoning polygon carrying a link to its governing ordinance text, so a click on the map reaches the rule that applies to it |
* Planned capabilities are designed and specified but not shipping. Nothing on this page marked shipped depends on them.
This is where the opportunity layers come from
The pipeline is the upstream half of the two platforms you have already seen.
Into Land Intelligence
The manifest and its layers load straight into the map — the opportunity layers panel, the data grid and the AI Analyst are all reading what came out of stage five.
Solution detail →Into the Execution Layer
Scored parcels become marketplace inventory; recertification urgency tiers and violation escalation tiers carry through to deadline monitoring with the same meaning on both sides.
Solution detail →Alongside parcel federation
The proxy answers parcel questions live on one URL. The pipeline runs the deep analysis in batch and ships files. Different jobs, complementary rather than competing.
Solution detail →Name a city and we will run it
Tell us the jurisdiction you work in. We will look at what it publishes, tell you honestly which of the fifteen analyses its open data can support, and demo the ones it can. The pipeline is available as source code, so your team can onboard the rest of your markets without us.