◆ Solution 04 · Data pipeline & jurisdiction onboarding

Onboard a city by writing a config file.

Municipal GIS is a thousand different schemas wearing the same shape. This pipeline reads a jurisdiction's endpoints out of a YAML file, downloads them, translates every column into a fixed canonical vocabulary, and runs the same spatial analysis it runs everywhere else — producing the opportunity layers that Land Intelligence and the Execution Layer consume.

🧠 Canonical semantic layer 🧮 DuckDB spatial SQL 📐 15 analytical queries 📋 YAML per jurisdiction
The problem

Every city is a rewrite

The analysis a developer wants is the same in every market: what is underbuilt, what is vacant, what is distressed, what can be assembled. The data underneath it never is.

🔤

Columns never match

Miami calls it FOLIO. Austin calls it PROP_ID. NYC calls it a BBL. Hard-code one and the query is worth nothing in the next county.

🕳️

Coverage is uneven

One city publishes liens and recertification records; the next publishes neither. A pipeline that requires a complete dataset never runs anywhere.

🐌

Extraction is slow and fragile

Paginated REST endpoints time out, cap result counts, and disagree about whether they support offsets at all. Half-written files corrupt the next stage.

🧳

The stack gets heavy

The obvious answer is a geopandas / GDAL / PostGIS deployment for something that should be a scheduled job producing a folder of files.

The fix

Five stages, driven entirely by the config file

The YAML names the endpoints, the field mappings and the outputs. Nothing else in the system changes when a jurisdiction is added.

STAGE 01 · EXTRACTESRI REST → GeoJSON

POST-based paginated download with offset or OID-range strategy auto-detected per layer, a thread pool across layers, retries, and atomic temp-then-rename writes so a failure never leaves a half-file behind.

STAGE 02 · CONVERT→ Hilbert GeoParquet

DuckDB sorts features along an ST_Hilbert space-filling curve and writes GeoParquet 1.1 with a bbox struct and ZSTD compression, so spatially local reads touch a handful of row groups.

STAGE 03 · CANONICALIZE→ semantic views

One DuckDB view per semantic role, generated from the config's field map. The parquet keeps its native column names; the view does the translation at query time.

STAGE 04 · ANALYZE→ opportunity layers

Fifteen city-agnostic spatial SQL queries run against the canonical views. Each writes a GeoJSON layer carrying source URLs and attribution on every feature.

STAGE 05 · MANIFEST→ MapLibre contract

A manifest describing each produced layer — title, geometry, styling, legend metric, exposed properties and upstream attribution — which the map application reads to build its UI.

◆ The core idea

One vocabulary, translated per city

Every analytical query in the system references canonical column names only — parcel_id, lot_size, zone_code, far. A jurisdiction's config maps those names onto whatever the county actually called them, and the pipeline emits a view that does the aliasing and the type coercion in one place.

  • Native data is never rewritten. The parquet keeps the county's original column names, so provenance survives all the way to the output.
  • Views are virtual — no duplication, no second copy of a 595,000-row parcel table.
  • Type coercion lives in the view, not scattered through the queries, and uses a non-throwing cast so a stray null in a source column cannot take a query down.
  • Adding a city adds a mapping. The analytical SQL is not touched, and a test asserts that no query can smuggle a city-specific column name back in.
DuckDB spatialGeoParquet 1.1 Python 3.10+pydanticYAML
the config says what things are called
# jurisdictions/miami_fl.yaml
sources:
  parcels:
    url: https://gisweb.miamidade.gov/…/MapServer/26
    layer_name: mdc_parcels
    srid: 2236
    attribution: "Miami-Dade County Property Appraiser"
    field_map:            # canonical: native
      parcel_id:  FOLIO
      address:    TRUE_SITE_ADDR
      owner:      TRUE_OWNER1
      lot_size:   LOT_SIZE
      land_value: LAND_VAL_CUR
the pipeline generates the translation
-- emitted automatically for role 'parcels'
CREATE OR REPLACE VIEW parcels AS
SELECT
  "FOLIO"                            AS parcel_id,
  "TRUE_SITE_ADDR"                   AS address,
  "TRUE_OWNER1"                      AS owner,
  TRY_CAST("LOT_SIZE" AS DOUBLE)     AS lot_size,
  TRY_CAST("LAND_VAL_CUR" AS DOUBLE) AS land_value,
  NULL                               AS building_count,
  geom, source_url, attribution
FROM mdc_parcels;
Unmapped means null, not broken. If a jurisdiction does not publish a canonical column, the view emits NULL for it rather than failing. Queries that depend on it degrade; every other query still runs. Partial configs are valid configs — which is what makes onboarding a city with thin open data possible at all.
Semantic roles

Thirteen roles a jurisdiction can fill

A config binds each role to one or more real layers. No city has all thirteen — and none needs to. Query availability is derived from which roles a jurisdiction actually bound.

RoleGeometryWhat it feeds
parcelsPolygonUnderbuilt, vacant, rezoning, assemblage, transit proximity, composite scoring
zoningPolygonUnderbuilt, rezoning, adaptive reuse, high-density near transit
zoning_overlaysPolygonAdaptive reuse
future_land_usePolygonRezoning candidates
building_footprintsPolygonUnderbuilt (FAR utilisation), tallest buildings
permitsPointStalled permits, active new construction, composite scoring
violationsPointOpen violations, composite scoring
liensPointLien hotspots, composite scoring
recertificationPointRecertification leads and clusters, composite scoring
transit_stationsPointTransit proximity, high-density zoning near transit
transit_linesPolylineReference context
water_servicePolygonVacant developable land
sewer_servicePolygonVacant developable land

* A single role can merge several upstream layers into one view — Miami's transit role, for example, unions the Metrorail and Metromover station services and tags each with its system name.

Analytical queries

Fifteen analyses, written once

Each produces a GeoJSON layer with its own property set, source URLs and upstream attribution. These are the opportunity layers that appear in Land Intelligence and the Execution Layer's marketplace.

LayerRoles it needsWhat it finds
Underbuilt parcelsparcels, zoning, footprintsLarge lots in non-low-rise zones built out to well under the floor-area ratio their zoning already permits
Vacant developable landparcels, water, sewerVacant parcels that already sit inside both a water and a sewer service area — the cheapest ground-up sites
Rezoning candidatesparcels, zoning, future land useParcels whose future land use designation permits far more density than the current zoning does
Assemblage candidatesparcelsAdjacent parcels under different ownership that combine into a materially larger development footprint
Adaptive reuse candidatesparcels, zoningBuildings whose current use no longer matches what the district is zoned for
Transit proximity parcelsparcels, transit stationsParcels inside a configurable walk-distance buffer of rapid transit — the TOD opportunity zone
High-density zoning near transitzoning, transit stationsDistricts that already allow the tallest, densest development and sit on transit
Recertification leadsrecertificationBuildings approaching or past their 40/50-year recertification, tiered by urgency against the current year
Recertification clustersrecertificationMulti-building addresses facing several recertification deadlines at once
Open violationsviolationsActive code violations, ranked by escalation tier — notice, lien, then lien plus overdue recertification
Lien hotspotsliens, parcelsParcels accumulating multiple compliance liens — an off-market distress signal
Stalled permitspermitsPermits still open long after issue, with the issue date exposed so the map can calculate staleness live
Active new constructionpermitsLive new-build, demolition and major-alteration permits — the current development pipeline
Tallest buildingsbuilding footprintsA reference benchmark set by reported height
Composite opportunity scoreparcels adaptiveA weighted multi-signal score over violations, liens, pending recertification and active permits

* The composite score adapts its own SQL to the roles a jurisdiction bound — a city without a liens layer still gets a score, computed from the signals it does publish. Owner and address exclusion filters are configurable per query, so public conservation land, parks, rights of way and utility plots can be kept out of the results.

Onboarding

Find the endpoints, map the fields, run it

No code change, no deployment, no schema migration — a new file and a command.

jurisdictions/austin_tx.yaml
schema_version: 1

jurisdiction:
  name: "City of Austin"
  state: TX
  county: "Travis"
  centroid: { lat: 30.2672, lon: -97.7431 }

sources:
  parcels:
    url: https://services.arcgis.com/…/FeatureServer/0
    layer_name: austin_parcels
    srid: 2277          # reprojected to 4326 on the way in
    field_map:
      parcel_id: PROP_ID
      address:   SITE_ADDR
      owner:     OWNER_NAME
      lot_size:  SQ_FT

  zoning:
    url: https://services.arcgis.com/…/FeatureServer/5
    layer_name: austin_zoning
    field_map:
      zone_code:  ZONING_TYPE
      far:        FAR_MAX
      max_height: HEIGHT_MAX

queries:                   # optional tuning per city
  underbuilt_parcels: { min_lot_sqft: 10000 }
  transit_proximity_parcels: { buffer_meters: 800 }
run it
# the whole pipeline
python -m opportunity_pipeline.run \
    jurisdictions/austin_tx.yaml

# or one stage at a time
python -m opportunity_pipeline.run \
    jurisdictions/austin_tx.yaml --stages extract,convert

# re-run only the analysis against parquet already on disk
python -m opportunity_pipeline.run \
    jurisdictions/austin_tx.yaml --stages analyze

# what lands on disk
data/austin_tx/
├── geojson/        # raw extraction, native columns
├── geoparquet/     # Hilbert-sorted, ZSTD, bbox struct
├── opportunities/  # the analysis layers
│   ├── 15_underbuilt_parcels.geojson
│   ├── 16_assemblage_candidates.geojson
│   └── 17_composite_opportunity.geojson
└── manifest.json   # the map application's contract
  • Queries skip, they do not crash. Every analysis declares the roles it requires; the engine checks what a jurisdiction bound and quietly skips the rest.
  • Analysis is resumable. Existing outputs are left alone unless forced, so regenerating one layer means deleting one file and re-running the stage.
  • Extraction is re-entrant. Files are written to a temporary name and renamed on completion, so an interrupted run never poisons the next one.
  • Configs are validated before anything downloads — endpoints, field mappings and required columns are checked, and the result is a report you can read rather than a stack trace.
manifest.json — what the map reads
{
  "jurisdiction": "miami_fl",
  "generated_at": "2026-08-04T06:00:00Z",
  "layers": [
    {
      "id": "underbuilt_parcels",
      "title": "Underbuilt Parcels",
      "type": "geojson",
      "url": "opportunities/15_underbuilt_parcels.geojson",
      "geometry": "polygon",
      "style": { "fill-color": "#e8a33d" },
      "legend": {
        "metric": "far_ratio",
        "label": "Built FAR / Allowed FAR"
      },
      "properties": ["parcel_id", "address",
                     "zone_code", "far_ratio"]
    }
  ],
  "attribution": ["Miami-Dade County Property Appraiser",
                   "City of Miami — Miami 21"]
}
Proving portability

Three markets, three completely different data regimes

The point of the abstraction is that these three share no column names, no service conventions and — in New York's case — not even the same kind of endpoint.

🌴

Miami, FL

The reference jurisdiction, and the deepest coverage — parcels, Miami 21 zoning, future land use, building footprints, permits, code violations, liens, 40-year recertification, Metrorail and Metromover, and water and sewer service areas.

  • All fifteen analyses available
  • Miami-Dade County + City of Miami GIS
⛵

Fort Lauderdale, FL

A second Florida city on an entirely separate GIS server with its own naming conventions — tax parcels, zoning districts, proposed land use, building footprints, permits and service requests standing in for the violations role.

  • Onboarded by config alone
  • Broward County + City of Fort Lauderdale GIS
🗽

New York, NY

The hard case. MapPLUTO and zoning districts come from ArcGIS; DOB permits, DOB and HPD violations and OATH ECB liens come from the city's open-data API instead — different transport, same canonical vocabulary downstream.

  • Mixed ESRI and open-data sources
  • NYC DCP, DOB, HPD, DoITT, MTA

* Coverage per city is bounded by what that jurisdiction publishes — Fort Lauderdale, for example, produces no lien or recertification layers because no such service exists to bind. Ask which analyses are available for the markets you care about.

◆ The text track

Half of zoning is prose, so it gets its own pipeline

Setbacks, permitted uses, dimensional standards and parking minimums live in ordinance text, not in a feature service. A companion pipeline discovers a city's code portal, pulls the zoning chapters, and stores them in a form the map and the AI Analyst can actually query — which is what backs ordinance search in Land Intelligence.

  • Portal discovery — Municode, American Legal Publishing, eCode360 and Qcode are probed by URL pattern first, at no cost, with web search and a language model only as fallback.
  • Chapter navigation — the table of contents is read to find the zoning chapters, with a keyword-only mode that skips model calls entirely.
  • Clean Markdown output per chapter, plus optional structured extraction into zone districts, permitted uses, dimensional standards, parking requirements and overlay districts.
  • Search-ready storage — a full-text index with stemming and BM25 ranking, plus overlapping content chunks sized for retrieval, so an answer can be cited back to the section it came from.
MunicodeAmerican Legal eCode360QcodeSQLite FTS5
◆ Politeness is a feature

It behaves on other people's servers

Municipal infrastructure is not built for scraping, and a pipeline that gets an organisation blocked is worse than no pipeline.

  • Concurrency is capped globally, with a deliberate delay after every successful fetch.
  • Batch runs are sequential across cities by design — parallelising them is a rate-limit posture decision, not an optimisation.
  • Results are cached, so a re-run costs nothing until you explicitly force a refresh.
  • Pagination is followed with a hard cap, so a badly behaved portal cannot turn one city into an unbounded crawl.
On scraping rights. Zoning codes are public law, but portals carry their own terms of use. Sources are recorded per document so the provenance of anything reaching your analysts is auditable, and we will review the portals covering your jurisdictions with you before anything is run at scale.
5
Pipeline stages
13
Semantic roles
15
Analytical queries
3
Jurisdictions onboarded
4
Runtime dependencies
YAML
Per new city

* Four runtime dependencies: a spatial SQL engine, an HTTP client, a config validator and a YAML parser. No geopandas, no fiona, no shapely, no GDAL install and no spatial database to operate — deliberately, so the pipeline runs as a scheduled job on ordinary infrastructure.

Where it is going

Shipped, and what is next

Stated plainly, because the difference matters when you are deciding what to build on.

CapabilityStatusWhat it means
Config-driven pipeline, canonical layer, 15 analyses Shipped Running against live municipal endpoints with a regression suite over the canonical SQL and the reference jurisdiction
Onboarding a city by hand Shipped Proven twice beyond the reference city, including one on a completely different class of endpoint
Assisted discovery Planned In design Crawl a city's service directory, classify each layer into a semantic role and draft the field mapping for a human to review — the crawl deterministic, the judgement calls assisted
Scheduled orchestration Planned In design Nightly refreshes per jurisdiction with caching keyed on upstream change, retries and run observability
Vector tile output Planned In design Automatic switch to single-file vector tiles when a layer outgrows what GeoJSON should carry, with the manifest declaring which form a layer took
Text and geospatial tracks joined Planned In design Each zoning polygon carrying a link to its governing ordinance text, so a click on the map reaches the rule that applies to it

* Planned capabilities are designed and specified but not shipping. Nothing on this page marked shipped depends on them.

◆ Source-code SDK for every solution

Name a city and we will run it

Tell us the jurisdiction you work in. We will look at what it publishes, tell you honestly which of the fifteen analyses its open data can support, and demo the ones it can. The pipeline is available as source code, so your team can onboard the rest of your markets without us.