Behavioral Claim Velocity Attribution
Generated 2026-07-22 from current behavioral_scores.geoparquet (83,008 tracts)
What is this report? When a physical hazard (hail, wind) damages a neighborhood, the actual insurance claims filed often differ from what the damage alone would predict. This analysis measures that difference — called claim velocity — and traces it back to specific neighborhood characteristics.
Baseline = 1.0×. A neighborhood at 1.0× files exactly as many claims as the physical damage model predicts. A neighborhood at 1.5× files 50% more claims than the damage alone would suggest. A neighborhood at 0.7× files 30% fewer.
Velocity is clipped at 3.5× to prevent extreme outliers from distorting pricing. The model is calibrated on Texas homeowner wind and hail claims (2018–2023) from the TX Department of Insurance.
Scale of the model: 19 data-ingestion pipeline stages (Census, CFPB, HMDA, Zillow, court records, state insurance-department complaints, and more) feed roughly 95 distinct demographic/behavioral signal columns. Those signals are combined into 4 composite indexes (defined in the glossary below), which in turn drive a 2-part statistical model (frequency + severity) with 4 fitted coefficients.
Four neighborhood types (cohorts) are identified by unsupervised clustering on those composite indexes plus the ~90 other underlying signals:
Stable ·
Transitional ·
Stressed ·
Assertive
TX TDI calibrated multipliers: Assertive 1.344× · Stressed 1.310× · Transitional 1.031× · Stable 0.865×
Glossary — what each term means
Dispute Culture Index — how likely residents are to formally push back on an outcome rather than
accept it: CFPB complaints, state insurance-department complaints, court filings, and dispatch disputes.
Higher = more likely to escalate a claim. Raises claim frequency (β=+0.836).
Claiming Culture Index — density of claim-adjacent businesses (public adjusters, roofing
contractors, chiropractors, diagnostic imaging) net of stabilizing institutions (churches, credit unions,
home-improvement stores). Its model effect is negative (β=−0.598) — a known multicollinearity
offset against Dispute Culture and Attorney Density, not a standalone "claims go down" signal.
Attorney & Adjuster Density Index — concentration of law offices and public adjusters serving
the neighborhood, a proxy for how likely a claim is to become represented/litigated.
Raises claim frequency (β=+0.688).
Maintenance Index — how well-kept and stable the housing stock is: owner-occupancy, long-tenure
residents, home-improvement loan activity, permit activity. The single strongest driver in the model —
well-maintained neighborhoods file far fewer claims than damage alone predicts (β=−1.216).
Cohorts (Stable / Transitional / Stressed / Assertive) — four neighborhood types found by
unsupervised clustering on the above indexes plus ~90 other demographic/behavioral signals.
Stable = high maintenance,
low dispute culture, lowest claim velocity.
Assertive = the opposite —
high dispute culture and attorney density, lower maintenance, highest claim velocity.
Transitional and
Stressed fall in between,
in that order.
Forecast (Part 4, below) — separate from everything above. Forecast predicts whether a tract's
cohort itself is likely to shift in the next 2-3 years (e.g. Transitional → Assertive), with a
label ("Unlikely" → "Will"), a predicted destination cohort, a horizon year, and a confidence score.
It is exploratory and not yet validated against outcomes — read it as a directional early-warning signal,
not a calibrated probability.
Velocity color key:
Below 0.90× — below baseline (fewer claims than modeled damage)
0.90×–1.10× — near baseline
1.10×–1.50× — elevated
Above 1.50× — high
Part 1A — How Each Index Moves Claim Velocity by Cohort
The model uses four composite neighborhood indexes, each built from dozens of raw signals. This table shows how much each index raises or lowers claim velocity for each cohort — relative to the median-level neighborhood (which anchors at 1.0×). Hover over any signal name for a full explanation. Click any column header to sort.
Claim Frequency Effect (how often claims are filed)
⚠️ Note on Claiming Culture Index: Its model coefficient is negative (β=−0.598) despite a positive weight in the Dispute Culture composite. This is a known multicollinearity artifact — the model uses it as a corrective offset. See the Gemini analysis for details.
Claim Severity Effect (how large individual claims are)
Part 1B — Individual Signal Contributions (Chain Estimate)
Each composite index is built from raw neighborhood signals. This table traces velocity contributions down to individual signals via chain propagation: a signal's effect = its weight in the parent index × the parent index's model coefficient. These are approximate estimates — not directly calibrated.
Source: Dallas–Fort Worth (1,559 census tracts). Hover signal names for definitions.
Part 2 — Velocity Distribution Across All Tracts
Every census tract in the national dataset is run through the model to get a predicted velocity. This section shows how those velocities are distributed within each cohort — not just the average, but the full spread from the 5th to 95th percentile. The 90% confidence interval on the average is computed by resampling 1,000 times (bootstrap).
| Cohort | Average Velocity | 90% Confidence Range | Tract Range (P5–P95) | Median Tract | Tracts in Cohort |
|---|
| Stable | 0.924× | 0.916× – 0.932× | 0.500× – 3.034× | 0.500× | 27,511 |
| Transitional | 1.424× | 1.411× – 1.435× | 0.500× – 3.500× | 0.777× | 26,257 |
| Stressed | 2.575× | 2.562× – 2.587× | 0.631× – 3.500× | 3.310× | 17,847 |
| Assertive | 3.290× | 3.282× – 3.299× | 1.746× – 3.500× | 3.500× | 11,393 |
Assertive cohort note: At the new ceiling of 3.5×, Assertive tracts may still saturate — the GLM predicts very high velocities for this cohort before any clipping. The TX TDI calibrated mean for Assertive is 1.344×, which is the empirically validated anchor. Individual-tract GLM predictions represent theoretical maximums, not observed averages.
Part 3 — Pricing Leverage by Signal
For each signal, all other signals are held at their median value while this one moves from low to extreme. This shows the standalone velocity curve for each signal — which one has the most pricing leverage, and where the non-linearity kicks in.
Pricing Leverage is Q95 velocity ÷ Q50 (median) velocity. A signal with leverage of 3.0× means an extreme-level neighborhood is priced at triple the velocity of an average one, based on this signal alone. The ⌈ marker indicates the model ceiling (3.5×) was hit — the true uncapped value is higher.
Protective signals (negative β) appear with leverage below 1.0× — meaning the Q95 tract has lower velocity than the median. Maintenance Index (β=−1.216) is the single strongest driver in the model, but its leverage ratio is inverted: a neighborhood at the 95th percentile for maintenance behaves far better than average, not worse. The leverage floor is just as valuable as the ceiling of risk signals.
Deferred: Multi-signal combination analysis (grid search, co-amplification) requires all 4 GLM predictors in the national dataset.
Claiming Culture Index and Attorney & Adjuster Density are currently absent from the national geoparquet — they are only available in city-level feature matrices.
Once added, the combination analysis will show whether high dispute culture and high attorney density together compound risk non-linearly.
Part 4 — Neighborhood Forecast: Which Tracts Are Predicted to Shift Cohort
Everything above measures a tract's current claim velocity. This section is different: it's the
model's Step-18 output predicting whether a tract's behavioral cohort itself is likely to change
in the next few years — e.g. a Transitional tract drifting toward Assertive as dispute activity rises and
maintenance investment falls. This is a directional early-warning signal, not a claim-velocity number, and
it has not been validated against real outcomes — treat the confidence scores as relative ranking,
not calibrated probabilities.
Aggregated across 198 of the model's city/state configs
(83,008 tracts — a handful of configs don't yet have forecast output).
Where "Likely"/"Will" tracts are headed
Current cohort × forecast label (% of each cohort)
Most common drivers cited for a predicted shift
- growing poverty share + rising eviction rate (8633 tracts)
- rising eviction rate + rising dispute activity (925 tracts)
- growing poverty share + rising dispute activity (440 tracts)
- rising complaint volume + growing poverty share (98 tracts)
- rising eviction rate + growing poverty share (96 tracts)
- rising complaint volume + rising eviction rate (59 tracts)
- rising complaint volume + rising dispute activity (58 tracts)
- growing poverty share + rising complaint volume (42 tracts)
Reading this: the large majority of tracts are forecast "Unchanged" — cohort shifts are the
exception, not the rule, over a 2-3 year horizon. Mean confidence sits well below 100 even for "Will"
tracts, reflecting that this is an exploratory signal built from short-window trend deltas (permit
activity, dispute-index momentum, price acceleration), not a validated predictive model.