Classification mix over time
The documented record
The Inspector General's July 2026 process inspection (OIG No. 26-E-03-FA0, July 28, 2026) sits alongside three other documents. A House Oversight Committee interim report (December 2025) concluded the then-chief pressured commanders to suppress crime statistics; a Department of Justice review reached similar conclusions the same month (coverage); and MPD's own internal affairs investigation (Final Investigative Report IS #26-000051, March 5, 2026, 554 pages; WJLA coverage) documented officials reclassifying violent assaults, robberies, and thefts under pressure from department leadership. The report issues findings on 18 officials, sustaining misconduct allegations against most of them, and states that all involved members were in a full duty status as of its March 5, 2026 date (p. 553); the subsequent discipline figures widely reported (19 officials charged, 13 facing termination) come from press coverage, not from the report itself. Federal prosecutors declined criminal charges (p. 20). The chief resigned in December 2025. The Washington Post's coverage frames the OIG report as the process audit explaining how that was possible. So this page asks two separately scoped questions: did the 2013 loss of oversight leave a mark, and did the documented 2023–2025 manipulation move the published totals.
This data sits inside a live dispute between federal and local officials over DC crime. This analysis does not adjudicate that dispute; it reports what the published data can and cannot show, with the confounds labeled.
Question 1: Did the 2013 loss of independent review shift the mix?
The question: When MPD's classification controls broke down in 2013, did the mix of offense categories shift?
The answer: Not attributably. 6 of 12 series do shift at the 2013 boundary after multiple-testing correction, but the identical test at placebo dates where nothing control-related happened fires comparably (up to 7 of 12). The classification mix shifts often throughout the archive, so the 2013 boundary cannot be singled out as a review-unit fingerprint. No evidence of a distinctive 2013 break is the result, and it is published as exactly that.
In July 2026 the DC Inspector General found that MPD disbanded its Staff Review Unit, the independent function that reviewed crime reports before they became official statistics, around 2013, and never replaced it (OIG No. 26-E-03-FA0, July 28, 2026). This page is different in kind from the eight disparity analyses: it audits the reliability of the incident archive itself. Our window extends earlier than the OIG's review period (2015–2025) because detecting a 2013 break requires data from before it. 6 of 12 series show a level shift at the January 2013 boundary that survives Newey-West errors and multiple-testing correction: theft (other) (+8.8 pp, 95% CI +2.1 pp to +15.5 pp); theft from auto (-8.4 pp, 95% CI -14.7 pp to -2.2 pp); assault w/ dangerous weapon (+0.6 pp, 95% CI +0.1 pp to +1.1 pp); theft from auto, as a share of all theft (-16.6 pp, 95% CI -28.7 pp to -4.6 pp); robbery, as a share of robbery + theft (-5.7 pp, 95% CI -9.4 pp to -2.0 pp); ADW, as a share of all violent categories (+2.8 pp, 95% CI +0.5 pp to +5.2 pp). But specificity fails: the identical test at placebo boundaries fires just as readily (up to 7 of 12 at dates where nothing control-related happened), so the honest reading is that the classification mix is generally nonstationary and the 2013 boundary does not stand out from that churn. This dataset provides no evidence of a distinctive review-unit fingerprint, and that null is the finding on the 2013 oversight question; the documented 2023–2025 manipulation era is tested separately below. This is a test of the published citywide totals, not an exoneration: case-level misclassification, and downgrades that move an incident out of the index categories entirely, would not appear here. The internal affairs report documents exactly that kind of case-level manipulation for 2023–2025, which is why Question 2 below tests that era directly.
6 of 12
Series (nine offense-category shares and three substitution pairs) with a level shift at January 2013 that survives Newey-West standard errors and Benjamini-Hochberg correction across the whole family. Read it with the placebo check below: a count here only means something if placebo dates fire less.
Each row: the category's share of that month's incidents (blue) and, for the nine categories, the raw monthly count (gray, each on its own scale). The counts are shown because the nine shares sum to 1: a real surge in one category mechanically depresses every other share, so a share break can be an echo of a real change elsewhere. Yellow ticks mark detected breaks. The five dashed reference lines and the two shaded bands:
- 2013: MPD disbands the Staff Review Unit, the independent function that reviewed crime reports before they became official statistics (per the OIG; no replacement until 2026).
- Aug 2015: the Mark43 records management system goes live at MPD (migration documented by the OIG; go-live date per the vendor). System migrations can shift recorded classifications on their own.
- Aug 2023: Pamela Smith becomes chief; the House Oversight and internal affairs reports date the documented pressure on crime classifications to her tenure.
- Dec 2025: the House Oversight and DOJ reports document crime-data manipulation and the chief resigns.
- Feb 2026: MPD launches mandatory DC Code classification training, the first of the post-OIG reforms; post-reform data is still thin.
- COVID-19 disruption: reported crime shifted abruptly for reasons unrelated to classification practice.
- Federal law enforcement surge begins Aug 2025: real enforcement, and possibly real crime, changed; nothing after this boundary is attributed to classification practice.
The a-priori test: a level shift at January 2013
One fixed break point, chosen before looking, at the OIG-documented loss of independent review. Each series gets a segmented interrupted time-series regression (level shift AND slope change at January 2013, calendar-month dummies), estimated on a symmetric window of five years each side of the break (2008–2017), with Newey-West standard errors (4 lags) because monthly shares are autocorrelated and naive OLS intervals would be too narrow. The symmetric window is load-bearing twice over: it stops long-run trend curvature from loading onto the break term, and it ends before the 2020 COVID discontinuity and the 2023 motor vehicle theft wave, so neither can masquerade as a 2013 effect.
| Series | Level shift at Jan 2013 | 95% CI (Newey-West) | p (NW) | p (BH-adjusted) | Survives correction |
|---|---|---|---|---|---|
| Theft (other) | +8.8 pp | +2.1 pp to +15.5 pp | 0.010 | 0.031 | Yes |
| Theft from auto | -8.4 pp | -14.7 pp to -2.2 pp | 0.008 | 0.031 | Yes |
| Motor vehicle theft | +0.9 pp | +0.0 pp to +1.8 pp | 0.045 | 0.078 | No |
| Robbery | -0.7 pp | -1.7 pp to +0.2 pp | 0.130 | 0.195 | No |
| Burglary | -1.2 pp | -2.9 pp to +0.5 pp | 0.161 | 0.214 | No |
| Assault w/ dangerous weapon | +0.6 pp | +0.1 pp to +1.1 pp | 0.016 | 0.032 | Yes |
| Sex abuse | +0.1 pp | -0.1 pp to +0.3 pp | 0.253 | 0.287 | No |
| Homicide | +0.0 pp | -0.0 pp to +0.1 pp | 0.263 | 0.287 | No |
| Arson | +0.0 pp | -0.0 pp to +0.0 pp | 0.515 | 0.515 | No |
| Theft from auto, as a share of all theft | -16.6 pp | -28.7 pp to -4.6 pp | 0.007 | 0.031 | Yes |
| Robbery, as a share of robbery + theft | -5.7 pp | -9.4 pp to -2.0 pp | 0.002 | 0.028 | Yes |
| ADW, as a share of all violent categories | +2.8 pp | +0.5 pp to +5.2 pp | 0.016 | 0.032 | Yes |
This is a family of 12 tests. At the 5% level, 12 uncorrected tests would be expected to produce about 0.6 false positives by chance alone, so no single uncorrected p-value below is headlined; the last column applies Benjamini-Hochberg across the family (q = 0.05). The three pair series use within-pair shares (for example, theft from auto as a share of all theft), which the compositional coupling between the nine global shares cannot move.
The specificity check: placebo break dates
The identical family test (same window shape, same errors, same correction) run at January boundaries where nothing control-related happened, all chosen so their windows stay clear of the COVID band. If placebo dates fire as often as 2013, the mix is churning generally and a 2013 count proves nothing about the review unit.
| Break date tested | Corrected discoveries |
|---|---|
| Jan 2013 (review unit disbanded) | 6 of 12 |
| Jan 2011 (placebo) | 1 of 12 |
| Jan 2012 (placebo) | 7 of 12 |
| Jan 2014 (placebo) | 4 of 12 |
| Jan 2015 (placebo) | 1 of 12 |
Placebo dates fire as often as the real boundary: the 2013 count carries no specificity.
Detected breaks, wherever they fall
Binary segmentation on each deseasonalized series, with no prior about where a break should be. A break landing on the Mark43 migration, in the COVID band, or in the federal-surge era rather than at a tested boundary is reported as exactly that.
| Series | Detected break month(s) |
|---|---|
| Theft (other) | Jun 2009; Feb 2012 (near: Review unit disbanded (2013)); Feb 2014; Feb 2017; Mar 2019 (near: COVID-19 band begins (Mar 2020)); Apr 2020 (COVID-era; not attributable to classification practice); Nov 2023 (near: Pressure era begins (Aug 2023)); Aug 2025 (federal-surge era; nothing past this boundary is attributed to classification practice) |
| Theft from auto | Jul 2009; Feb 2011; Nov 2013 (near: Review unit disbanded (2013)); Feb 2016 (near: Mark43 RMS goes live (Aug 2015)); May 2020 (COVID-era; not attributable to classification practice); Nov 2022 (near: Pressure era begins (Aug 2023)); Sep 2025 (federal-surge era; nothing past this boundary is attributed to classification practice) |
| Motor vehicle theft | Jan 2009; Feb 2010; Feb 2011; Feb 2012 (near: Review unit disbanded (2013)); Jul 2020 (COVID-era; not attributable to classification practice); Nov 2022 (near: Pressure era begins (Aug 2023)); Nov 2023 (near: Pressure era begins (Aug 2023)); Jul 2025 (at the federal surge boundary; not attributable to classification practice) |
| Robbery | Jan 2014 (near: Review unit disbanded (2013)); Oct 2016; Jun 2020 (COVID-era; not attributable to classification practice); Mar 2023 (near: Pressure era begins (Aug 2023)); Oct 2024 (near: Federal surge begins (Aug 2025)) |
| Burglary | Aug 2010; Nov 2011; Dec 2013 (near: Review unit disbanded (2013)); Jan 2015 (near: Mark43 RMS goes live (Aug 2015)); Oct 2016; Apr 2022 |
| Assault w/ dangerous weapon | May 2010; Sep 2016; May 2018; Apr 2020 (COVID-era; not attributable to classification practice); Dec 2021; Dec 2022 (near: Pressure era begins (Aug 2023)); Sep 2025 (federal-surge era; nothing past this boundary is attributed to classification practice) |
| Sex abuse | Dec 2011; Aug 2018; Jul 2024 (near: Pressure era begins (Aug 2023)) |
| Homicide | Dec 2010; Mar 2018; Apr 2020 (COVID-era; not attributable to classification practice); Mar 2025 (near: Federal surge begins (Aug 2025)) |
| Arson | Apr 2009; Jul 2013 (near: Review unit disbanded (2013)); Mar 2016 (near: Mark43 RMS goes live (Aug 2015)) |
| Theft from auto, as a share of all theft | Jun 2009; Feb 2017; Aug 2022 (near: Pressure era begins (Aug 2023)); Aug 2023 (near: Pressure era begins (Aug 2023)); Aug 2025 (federal-surge era; nothing past this boundary is attributed to classification practice) |
| Robbery, as a share of robbery + theft | Aug 2009; Mar 2012 (near: Review unit disbanded (2013)); Feb 2014; Jun 2016 (near: Mark43 RMS goes live (Aug 2015)); Apr 2020 (COVID-era; not attributable to classification practice); Feb 2023 (near: Pressure era begins (Aug 2023)); Oct 2024 (near: Federal surge begins (Aug 2025)) |
| ADW, as a share of all violent categories | Mar 2010; Jan 2014 (near: Review unit disbanded (2013)); Jan 2022; Feb 2023 (near: Pressure era begins (Aug 2023)); Apr 2024 (near: Pressure era begins (Aug 2023)); Jul 2025 (at the federal surge boundary; not attributable to classification practice) |
Question 2: Did the documented 2023–2025 manipulation move the published totals?
The question: During the era where manipulation is documented (2023–2025), did the published offense mix move the way the documented downgrade paths predict?
The answer: No aggregate signature. None of the four directional contrasts predicted by the documented downgrade paths survives correction. A null here does not mean manipulation did not happen: manipulation is documented at case level in the 554-page internal affairs report. It means the documented case-level manipulation was not large enough to visibly move citywide monthly aggregates in the public index-crime feed. Given the documented scale, a null is the expected result; the magnitude bridge below shows why.
Unlike 2013, this era comes with documented direction: the internal affairs report describes robberies logged as theft, weapons assaults downgraded below ADW, and thefts shifted to lesser categories. None of the four one-sided contrasts built from the internal affairs report's documented downgrade paths (robberies logged as theft, weapons assaults downgraded, thefts shifted to lesser categories) survives Benjamini-Hochberg correction at the August 2023 boundary. A null here does not mean manipulation did not happen: manipulation is documented at case level in the 554-page internal affairs report. It means the documented case-level manipulation was not large enough to visibly move citywide monthly aggregates in the public index-crime feed. The magnitude bridge below shows the documented scale falls below this test's detection floor, so this null is the expected result given the documented scale. The decline-then-reversion signature that aggregate-level manipulation would leave is ABSENT: no contrast both falls during the pressure era and reverts after exposure.
The era test reuses Question 1's machinery at the August 2023 boundary (the first full month of the tenure the House and internal affairs reports date the pressure to): segmented regression on a symmetric window of 24 months each side (Aug 2021–Jul 2025), which by construction avoids the COVID band and ends exactly at the August 2025 federal surge boundary. Newey-West errors throughout. Three families, each Benjamini-Hochberg corrected separately and labeled: the 12-series two-sided family, the 4-contrast one-sided directional family, and the 4-contrast reversion family.
The directional family (one-sided, signs fixed by the documented paths)
| Contrast | Documented path (predicts a fall) | Shift at Aug 2023 | p (one-sided, NW) | p (BH) | Fires |
|---|---|---|---|---|---|
| ADW, as a share of all violent categories | weapons assaults downgraded below ADW (many exit the feed) | -1.2 pp | 0.381 | 0.550 | No |
| Robbery, as a share of robbery + theft | robberies logged as theft | -0.7 pp | 0.413 | 0.550 | No |
| Theft from auto, as a share of all theft | thefts shifted toward the lesser theft category | -2.1 pp | 0.063 | 0.250 | No |
| Assault w/ dangerous weapon | ADW share of the whole mix falls as downgrades exit | +0.2 pp | 0.819 | 0.819 | No |
The reversion test (post-exposure step, Dec 2025)
If the declines were manipulation, exposure should end them: the December 2025 House and DOJ reports and the resignation are the exposure boundary. Two caveats travel with this table: only 9 post-exposure months exist, and every one of them sits past the August 2025 federal surge boundary, so reversion cannot be cleanly attributed to exposure versus the surge. A post-exposure step with no matching pressure-era decline is NOT the signature: both legs are required, and a lone step can as easily reflect the surge-era mix.
| Contrast | Post-exposure step (Dec 2025) | p (one-sided, NW) | p (BH) | Fires |
|---|---|---|---|---|
| ADW, as a share of all violent categories | +2.8 pp | 0.120 | 0.231 | No |
| Robbery, as a share of robbery + theft | +1.6 pp | 0.174 | 0.231 | No |
| Theft from auto, as a share of all theft | +0.3 pp | 0.442 | 0.442 | No |
| Assault w/ dangerous weapon | +1.5 pp | 0.001 | 0.005 | Yes |
The magnitude bridge: could this test even see it?
The primary document (Final Investigative Report IS #26-000051, March 5, 2026) states no comprehensive total and calls its fullest tally "by no means a complete record" (p. 526), so the total here is our computed sum of the report's own per-official audit counts: Cmdr M. Pulliam 157 (p. 16); Capt Donigian 363 (p. 318); Capt Haskis 194 (p. 334); Capt Rivers 131 (p. 217); Capt R. Pulliam 47 (p. 370); Capt Merrill 30 (p. 533). By that count, the internal affairs investigation documents 922 improperly edited or reclassified reports across six named officials, the large majority of them theft reports moved out of the nine public index categories; the documented robbery and ADW reclassifications number 5 and 11 respectively. Press coverage's named figures (Pulliam 157, Donigian about 360) match the primary audits. For scale: the pre-surge era window averaged roughly 185 robberies, 87 ADW incidents, and 1070 thefts reported citywide per month.
| Scenario | Implied effect | Minimum detectable effect (one-sided 5%) |
|---|---|---|
| Documented robbery path: 5 reports (pp. 16, 256, 309) | 0.014 pp | 5.25 pp on robbery share of robbery+theft |
| Documented ADW path: 11 reports (p. 16) | 0.093 pp | 6.44 pp on ADW share of violent |
| Documented theft exits: 406 reports counted (pp. 16, 217, 335, 370, 533), most of Donigian's 363 on top (p. 318) | 14/month leave the feed (0.6% of monthly volume) | Invisible by construction: exits never reach this dataset |
| Stress ceiling: all 922 edits as robbery downgrades | 2.62 pp | 5.25 pp |
| Stress ceiling: all 922 edits as ADW downgrades | 8.72 pp | 6.44 pp |
Documented-scale manipulation falls below what this test can detect, so a null is the expected result given the documented scale. Even the stress ceiling, all 922 edits concentrated on the robbery path, stays under that contrast's detection floor; only if nearly all of them had been ADW downgrades would the shift have been visible, and the report's documented ADW count is 11.
The era boundary and its placebos (two-sided, all 12 series)
| Series | Level shift at Aug 2023 | 95% CI (Newey-West) | p (NW) | p (BH-adjusted) | Survives correction |
|---|---|---|---|---|---|
| Theft (other) | +3.8 pp | +0.4 pp to +7.1 pp | 0.027 | 0.109 | No |
| Theft from auto | +0.2 pp | -3.0 pp to +3.4 pp | 0.897 | 0.897 | No |
| Motor vehicle theft | -5.1 pp | -8.5 pp to -1.7 pp | 0.003 | 0.042 | Yes |
| Robbery | +0.4 pp | -2.8 pp to +3.7 pp | 0.801 | 0.897 | No |
| Burglary | +0.2 pp | -0.5 pp to +0.9 pp | 0.615 | 0.897 | No |
| Assault w/ dangerous weapon | +0.2 pp | -0.3 pp to +0.7 pp | 0.362 | 0.725 | No |
| Sex abuse | +0.1 pp | -0.2 pp to +0.5 pp | 0.522 | 0.894 | No |
| Homicide | +0.2 pp | +0.0 pp to +0.3 pp | 0.026 | 0.109 | No |
| Arson | -0.0 pp | -0.0 pp to +0.0 pp | 0.077 | 0.231 | No |
| Theft from auto, as a share of all theft | -2.1 pp | -4.8 pp to +0.6 pp | 0.125 | 0.300 | No |
| Robbery, as a share of robbery + theft | -0.7 pp | -7.0 pp to +5.5 pp | 0.825 | 0.897 | No |
| ADW, as a share of all violent categories | -1.2 pp | -8.9 pp to +6.5 pp | 0.762 | 0.897 | No |
| Break date tested | Corrected discoveries |
|---|---|
| Aug 2023 (documented pressure era begins) | 1 of 12 |
| Aug 2011 (placebo) | 1 of 12 |
| Aug 2013 (placebo) | 0 of 12 |
| Aug 2015 (placebo) | 1 of 12 |
| Aug 2017 (placebo) | 0 of 12 |
Placebo dates fire as often as the real boundary: the Aug 2023 count carries no specificity.
Placebo limitation, stated: a window of this shape that avoids both the COVID band and the federal surge fits only the real boundary, so these placebos come from the pre-COVID era. They test the machinery's false-fire rate in calmer data, not a perfectly matched counterfactual.
The district-level test, and the trap it documents
The obvious next step after a citywide null is to de-dilute: about 21 documented edits a month vanish against thousands of citywide incidents, but concentrated into single districts the same edits loom much larger. The internal affairs report names the districts and the windows itself, and one cell looked large enough to see from outside: the Seventh District, January to October 2024, where the report's audit window and per-official counts put roughly 15 theft reports a month leaving the public feed against a base of about 75 (pp. 88, 217, 546). We built that test, pre-registered its interpretation before estimating (database/theft_exit.py, wording committed before the estimator existed), and ran it.
The design does not hold, and that is the finding. A difference-in-differences estimate is only evidence if the treated district moved in parallel with the comparison districts before the window; the Seventh District did not. Its theft counts rose 38% in 2023 (806 to 1,113), the steepest rise in the city that year, and then fell 16% in 2024. The comparison districts are not a quiet backdrop either: over the same year they ranged from -0.3% to +31%, so the pre-window period is one where districts were moving sharply and differently from one another. Weighing the whole twelve-month lead path against the same placebo assignments the estimate would use, the pre-registered permutation joint lead test rejects parallel pre-trends at p = 0.016 against a gate of 0.10. The one cell whose documented scale approached visibility therefore cannot be estimated, and the remaining cells sit below their detection floors. The pre-registered branch for this outcome (design-fails) says it plainly: the honest output is that answer, not a weaker substitute test.
Seventh District theft (theft/other plus theft from auto), monthly, as log-point deviations from its own late-2022 baseline relative to the comparison districts, with a 95% band. The audit window (p. 217) is shaded; the first THRIVE area in the district (May 2024, p. 13) is the dashed line, marked separately so the window-versus-initiative timing stays readable. The divergence is the 2023 lead-up, before either boundary.
Two further failures converge on the same verdict, one statistical and one supplied by the report itself. First, every documented cell sits below its placebo-derived detection floor, the Seventh District marginally so:
| Documented cell | Documented edits/mo | Base thefts/mo | Expected dip | Detection floor | Verdict |
|---|---|---|---|---|---|
| Seventh District, January-October 2024 | 15.1 | 75 | 16.7% | 18.4% | below detection floor |
| Second District, January-August 2025 | 15.4 | 358 | 4.1% | 14.4% | below detection floor |
| Fifth District (Rosedale THRIVE), February-August 2025 | 7.4 | 232 | 3.1% | 15.3% | below detection floor |
Detection floors here come from the placebo distribution itself (the deficit size placebo district-windows produce 5% of the time), which at seven districts is wider than asymptotic approximations suggest. A cell below its floor cannot confirm or rule out the documented pattern; its null is uninformative by construction.
Second, the report's own findings place commander-directed Seventh District misclassification across 2023 and 2024 (pp. 88, 541–542), which means the comparison baseline year is itself a documented editing year: the design has no clean pre-period available in that district for reasons the primary document supplies. The pre-trend verdicts:
| Cell | Permutation joint lead p | Placebo assignments | Gate (fails at p ≤ 0.10) |
|---|---|---|---|
| Seventh District, January-October 2024 | 0.016 | 63 | FAIL |
| Second District, January-August 2025 | 0.737 | 18 | pass |
| Fifth District (Rosedale THRIVE), February-August 2025 | 0.421 | 18 | pass |
Composed with the citywide result above, this is the substantive finding: at both resolutions available in public data, citywide aggregates and district-level contrasts, documented case-level manipulation leaves no testable trace. The manipulation is established by the internal affairs report's sustained findings, not by this page; what this page establishes is that the public feed could not have caught it, and that the one analysis that looks like it catches it is a trap. With no measurable deficit in the Seventh District cell, the window-versus-THRIVE timing comparison has nothing to separate, and no timing claim is made.
How this number is built (and where it's soft)
- The prompt for this page. The DC OIG's inspection of MPD's crime data reporting (OIG No. 26-E-03-FA0, July 28, 2026) found a prolonged breakdown of internal controls: the Staff Review Unit was disbanded around 2013 during a records-system transition and never replaced, approved reports could be edited with minimal oversight, and some classification policies had gone decades without update. The OIG's process review did not itself adjudicate intent; the separate House, DOJ, and internal affairs findings in the context block document case-level manipulation for 2023–2025. This page tests for visible statistical fingerprints; it does not assume or rule out either.
- Construction. Monthly shares of the citywide classification mix over 224 months (2008–2026, through the last complete month), computed live from the incident archive.
- Seasonality. Reported crime is strongly seasonal. The changepoint scan runs on deseasonalized series (calendar-month means removed); the 2013 regression handles it with calendar-month dummies instead.
- Detected breaks vs the a-priori test, kept separate. Binary segmentation (mean-shift cost, BIC-style penalty at 3σ² log n, 12-month minimum segments) scans the FULL window with no prior about location; the interrupted time-series test fixes January 2013 in advance. They answer different questions and are never pooled.
- The 2013 test is local by design. Segmented regression (level shift plus slope change) on 60 months each side of the break. Over the full 18-year window, a single-trend specification would let any long-run curvature in a series (and the compositional echo of the 2023 motor vehicle theft wave) load onto the break term and manufacture shifts; the symmetric local window is the standard defense, at the cost of saying nothing about slow drift that starts years after 2013.
- Serial correlation. Monthly shares are autocorrelated, so the 2013 regression uses Newey-West (Bartlett kernel, 4 lags) standard errors. Naive OLS p-values are computed for audit but never used for any conclusion.
- Specificity is tested, not assumed. A significant count at 2013 means nothing if the same machinery finds "breaks" everywhere, so the identical family test is run at placebo January boundaries (2011, 2012, 2014, 2015, each with a COVID-free window). The finding sentence at the top is driven by that comparison: when placebos fire as often as 2013, the page says the mix is churning generally and claims no fingerprint.
- Multiple testing. 12 series are tested, so the family-wise expectation is about 0.6 chance "significant" results at the 5% level. Discoveries are Benjamini-Hochberg corrected (q = 0.05); the page never headlines an uncorrected hit.
- Shares are compositional. The nine shares sum to 1, so a genuine surge in one category (a real motor-vehicle-theft wave, say) drags every other share down with zero classification change anywhere. Raw counts are drawn beside every share, and the substitution pairs use within-pair shares that the coupling cannot move.
- The 2020 discontinuity is COVID. Reported crime shifted abruptly in March 2020 for reasons unrelated to classification practice. The band is marked on the chart, breaks inside it are labeled COVID-era, nothing in that window is attributed to classification, and the a-priori 2013 test ends in 2017, before the band, by construction.
- Alternative explanations, listed not dismissed. A break in a share series can reflect real crime changes, reporting-rate changes, policy or legal changes, records-system migrations (the OIG dates the Cobalt-era RMS transition to around 2013, the same year the review unit was disbanded; the Mark43 go-live is dated August 2015 by the vendor), or classification practice. This page can date breaks; it cannot adjudicate between those causes.
- The era test (Question 2). Same segmented regression at the August 2023 boundary on a ±24-month window (Aug 2021–Jul 2025): the window avoids the COVID band and ends at the federal surge by construction. The directional family fixes signs a-priori from the internal affairs report's documented downgrade paths and tests one-sided; the reversion family tests the post-exposure step (Dec 2025) with era, surge, and exposure dummies in one regression. Each family is Benjamini-Hochberg corrected separately.
- The August 2025 federal surge is an attribution boundary. Real enforcement, and possibly real crime, changed when the surge began. It is shaded on the chart like COVID, detected breaks past it are labeled, the era window ends at it, and the reversion test's caveats name it. Nothing after August 2025 is attributed to classification practice.
- The district-level test (design and pre-registration). Stacked difference-in-differences event study on log monthly theft counts (theft/other plus theft from auto; motor vehicle theft excluded a-priori because the Kia/Hyundai wave would dominate any district contrast) for the three (district, window) cells the internal affairs report itself names, with comparison district-months as controls. All finding language, including the branch that fired, was committed to version control before the estimator existed; the estimator, tests, and this page's wording are in the repository as the audit trail.
- A-priori exclusions. Two families of tests were excluded before any estimate was computed: per-district robbery-share contrasts and per-district ADW-share contrasts. The internal affairs report's documented robbery-path and ADW-path reclassifications total roughly 5 and 11 reports respectively (pp. 16, 256, 309); at every district's denominator that scale sits one to two orders of magnitude below any detection floor. Running those tests could only produce uninformative nulls, and reporting uninformative nulls alongside real results would overstate what was checked. The exclusion, and the finding language for every branch below, were committed to version control before the estimator existed.
- Inference at seven districts. All p-values are permutation ranks in placebo ensembles (placebo districts and placebo window placements); cluster asymptotics are not trusted at n = 7. The pre-trend gate uses the same inference as the estimate itself. With seven districts, asymptotic joint tests have unknown calibration, and this gate can kill the entire analysis, so its p-value cannot rest on an approximation the primary test already rejects. The joint lead statistic is therefore recomputed under the same placebo procedure as the primary estimate: the treated label is reassigned across comparison districts and across placebo window placements, the statistic is recomputed for each assignment, and the gate's p-value is the rank of the real statistic in that placebo distribution.
- The comparison pool. The Third District is excluded from the comparison pool from June 2024 through the end of the estimation span. The report documents 157 reclassifications by the then-Third District commander (p. 16) but gives no audit date range; the six-month span comes from one witness's characterization (p. 71) and the May 12, 2025 discovery date (p. 15). Because the window is inferred, the exclusion is deliberately over-wide, roughly twice the characterized span plus margin, so that no plausibly contaminated Third District month anchors the comparison. The estimate is reported both with this exclusion (primary) and without it (sensitivity). On this run nothing fires either way: over the two pre-trend-passing cells the pooled coefficient is +0.088 log points (one-sided permutation p = 0.88; contribution weights 2D-2025 53%, 5D-2025 47%), and with the Third District included it is +0.096 (p = 0.92). Under the pre-registered rules a sensitivity run cannot upgrade a primary null, and neither of these is a deficit to begin with.
- THRIVE is a competing explanation. THRIVE, a targeted enforcement initiative, operated in the Fifth and Seventh District areas during these windows, and an initiative that genuinely reduced theft would leave the same district-shaped deficit as reports edited out of the feed. This data cannot separate those two readings; the event-study plots mark both the audit windows and the THRIVE spans so the timing can be compared, and the finding text carries the ambiguity rather than resolving it. On this run no deficit exists to compare against the initiative windows, so no timing claim is made.
- District data quality. Before any estimation, 1,564 archive rows missing a district assignment were recovered by point-in-polygon against the DCGIS police-district boundaries (336 of them a single November 2024 ingest batch); zero null-district rows remain. District boundaries are stable across the estimation span (block-level churn under 1% a year, all post-2012-realignment).
- What a null means, on either question. A null at 2013 is a finding about the visible mix on the oversight question only. A null on the era test does not mean manipulation did not happen (it is documented at case level in the internal affairs report); it means the documented case-level manipulation was not large enough to visibly move citywide monthly aggregates in the public index-crime feed, and the magnitude bridge quantifies whether it even could have. Neither null is proof classifications were accurate: downgrades out of the index set are invisible here by construction.
Related: Enforcement vs reported crime | Methodology | All analyses