◆ Solution 07 · Ordinance search & AI extraction

Ask the ordinance a question. Get the passage it came from.

Setbacks, height limits, FAR, parking minimums and permitted uses live in municipal legal text, not in a table. This turns that text into two things a system can actually use: natural-language answers that carry their own citations, and structured dimensional data extracted at scale with a full audit trail.

📚 27,926 ordinance nodes 🔍 BM25 retrieval 🧾 Cited answers 🧪 Audited AI outputs
The problem

The rules that decide a project are prose

Everything else about a parcel is a number in a database. What you are legally allowed to build on it is a chapter of municipal code with cross-references.

📖

It reads like law

Because it is. A single answer can depend on a district, an overlay, a use table, a definitions section and an exception three chapters away.

🏙️

Every city writes its own

Same concepts, different words, different structure, different numbering. Nothing transfers between municipalities.

🎲

A chatbot will guess

Ask a general-purpose model about a setback and it will answer confidently from nothing. In a regulatory context that is worse than no answer.

⏳

It does not scale by reading

Screening one site by hand is an afternoon. Screening a market is a headcount problem, and the answer goes stale when the code is amended.

Two approaches, one corpus

Answers for people, structure for systems

These are different jobs and they are built differently — deliberately. One answers a question in context; the other turns a whole municipality into rows.

💬

Retrieval-augmented Q&A

A question retrieves the most relevant ordinance passages first; the model is then constrained to answer from those passages alone, and the passages come back with the answer. Built for the moment someone needs to know whether a thing is allowed here.

  • Full-text retrieval with BM25 ranking over RAG-sized chunks
  • The response carries exactly the passages given to the model
  • Retrieval miss returns a "couldn't find" answer rather than an unaided guess
  • Answers exportable as JSON, CSV, Markdown or plain text
🗄️

In-database AI extraction

A SQL pipeline that ingests providers, runs model inference inside the database, and writes structured rows — dimensional standards per district, permitted-use classifications, and the full prompt and raw response beside each one.

  • Setbacks, height, FAR and density as typed columns per district
  • Use permissions classified as permitted, conditional, prohibited, accessory or unclear
  • Every prompt, raw response and parsed result retained as an audit trail
  • Exports to CSV, GeoJSON or PostGIS
20
South Florida sites
297
Code sections
27,926
Content nodes
22,395
Retrieval chunks
24,026
Resolved links
1,759
Extracted tables

* Figures describe the current South Florida corpus build. The corpus is a snapshot, not a live mirror — it is produced by a separate acquisition pipeline and refreshed by rebuilding, so a recently amended ordinance is only present after the next build. Ask which municipalities and which build date apply to your markets.

◆ Grounded answers

Every answer arrives with its evidence

Retrieval runs first and the model is constrained to what it returned. The sources array on the response is not a bibliography added afterwards — it is exactly the set of passages the model was given, which is what makes an answer checkable rather than merely plausible.

  • Auditable by construction — read the answer, then read the ordinance text it was built from, in the same response.
  • Refuses rather than invents — when retrieval finds nothing relevant, the endpoint says so instead of letting the model fill the gap.
  • Batch questions — up to ten at a time, with export, for screening a site against a checklist rather than one query at a time.
  • Direct text access too — section hierarchy, individual nodes and extracted tables are available without touching a model at all, and need no API key.
FastAPISQLite FTS5 BM25OpenAPI 3.1Pydantic
/docs — POST /ai/query
Swagger UI for the natural-language query endpoint, describing the retrieval-then-completion pipeline and explaining exactly what the confidence field does and does not mean.
The endpoint documents its own limitsIncluding a plain statement that the confidence value reflects only whether retrieval succeeded — and should not be thresholded on.
On that confidence value. It is fixed at 0.85 whenever any context was retrieved and 0.0 when none was. It tells you retrieval worked. It is not a calibrated measure of whether the answer is correct, and the API documentation says so at the point of use rather than in a footnote. We would rather you knew that than be impressed by a number.
◆ Structured extraction

Inference where the data already is

The extraction pipeline runs model inference from inside the analytical database, so the prompt, the raw response and the parsed result land in a table beside the structured output rather than disappearing into an application log.

  • Two provider tiers — ordinance-text providers for nuanced reasoning, structured data APIs for dimensional facts, ingested through the same schema.
  • Cross-provider validation — where two sources describe the same district, disagreements are recorded as flags with warning or error severity rather than silently resolved.
  • Reproducible by stage — setup, ingest, AI, report and validate run independently, so a rerun does not mean redoing everything.
  • Assembled reports — a structured feasibility report per parcel request, exportable to CSV, GeoJSON or PostGIS for whatever consumes it next.
DuckDBIn-DB inference Audit trailPostGIS export
run a district through the pipeline
# Ingest a municipality from an ordinance provider
python scripts/run_pipeline.py --provider gridics \
    --municipality-id 1 --doc-id 45

# Ask a reasoning question about one district
python scripts/run_pipeline.py --stage ai \
    --municipality-id 1 --zoning-district "T6-8-O" \
    --user-query "Is a hotel permitted only with a \
                 conditional use permit?"

# Everything: setup → ingest → ai → report → validate
python scripts/run_pipeline.py --stage all \
    --provider gridics --municipality-id 1

# What lands in the database
SELECT zoning_district, max_height_ft, max_far,
       front_setback_ft, min_lot_area_sf
FROM   dimensional_attributes
WHERE  municipality_id = 1;

# …and the prompt that produced each one
SELECT task_type, prompt_text, raw_response, parsed_json
FROM   ai_outputs
WHERE  task_type = 'extraction';
API surface · OpenAPI 3.1

Four groups, and one that needs no key

The text endpoints are just a database read — no model involved, so they work without an AI provider configured at all. Everything that costs tokens is grouped where you can see it.

GroupWhat it doesNeeds a model?
ZoningSites, section hierarchy, ordinance search, individual nodes and the tables extracted from themNo
AINatural-language questions, batch questions, scenario decisions, zone comparison, node summaries, generated SQLYes
AnalysisZoning standards per district, parcel upside, ranked opportunities, district statisticsNo see limits
PredictiveScenario value projections, property summaries, investment memosPartly
HealthLiveness, plus database and model connectivity reported independentlyNo
Swagger UI overview for the Zoning Real-Estate API, describing the data source, the prerequisites for each endpoint group, response formats and conventions.
Prerequisites up frontThe overview states which endpoint groups need a model key, which need a derived table, and which need live network access — before you hit one and get an error.
Swagger UI endpoint listing showing the Health, Zoning and AI route groups with their individual operations.
Text access is separate from AIThe Zoning group reads scraped ordinance text directly and is explicitly documented as working without an API key.
◆ Feasibility scenarios

What the zoning permits, against what is built

Where structured standards exist for a district, they can be compared with what a parcel actually is — producing utilisation ratios, headroom, and value projections across a set of named development scenarios with every assumption stated.

  • Four scenarios — hold as-is, renovate, resolve compliance, or redevelop to the zoning envelope.
  • Assumptions are inputs, not magic — appreciation rate, construction cost per square foot, soft costs, developer margin, cap rate and rent are all yours to set, and default values are visible.
  • Utilisation, not just capacity — current versus allowed FAR, density and height, with the additional floor area and units the envelope would permit.
  • Grounded narrative — generated summaries and memos are written from the computed metrics rather than invented alongside them.
/docs — POST /predictive/analyze
Swagger UI for the predictive analysis endpoint, listing the four development scenarios and the editable market assumptions in the request body.
Modelled estimates, not appraisalsThe endpoint group says exactly that in its own description, and every assumption behind a projection is an explicit input.
Before you rely on it

What this does not do

This is software that reads law and produces numbers. Being precise about where it is weakest is the only responsible way to sell it.

⚖️

It is not a legal opinion

Answers are a research aid that points you at the governing text. Verify against the ordinance before anything is filed, purchased or built. The citations exist precisely so that verification is quick.

📐

The standards table is noisy

The per-district standards behind the analysis endpoints are extracted from scraped ordinance tables by pattern matching, and known-bad values exist in the current build. Treat them as a starting point to confirm, not as authority.

🎚️

Confidence is not correctness

The confidence field reports whether retrieval found context — nothing more. It is documented that way in the API itself, and no downstream logic should threshold on it.

📸

The corpus is a snapshot

Ordinance text is captured by a separate acquisition run and refreshed by rebuilding. There is no incremental update, so an amendment appears only after the next build.

🌎

Coverage is where the corpus is

The current build is South Florida. Adding a market means running the acquisition pipeline against its code portal — routine, but it is a build step, not a configuration flag.

🔑

You bring the model

The AI endpoints need your own provider key, and several providers are supported. Text search, section hierarchy and table extraction run with no model and no key at all.

* We would rather show you this list on the way in than have you discover it in a deal. Generated SQL is gated so that only read statements execute, and the text endpoints are read-only against a snapshot — but none of that substitutes for checking a number against the ordinance that governs it.

◆ Source-code SDK for every solution

Bring us a question you already know the answer to

The fastest way to judge this is to ask it something you have already researched by hand, then read the passages it cites. Tell us the municipality and we will run your questions against it — and say plainly where the corpus does not yet reach.