Justice Lens

Classification mix over time

Published July 2026 · Data window: MPD reported incidents, 2008–present, monthly, citywide

The documented record

The Inspector General's July 2026 process inspection (OIG No. 26-E-03-FA0, July 28, 2026) sits alongside three other documents. A House Oversight Committee interim report (December 2025) concluded the then-chief pressured commanders to suppress crime statistics; a Department of Justice review reached similar conclusions the same month (coverage); and MPD's own internal affairs investigation (Final Investigative Report IS #26-000051, March 5, 2026, 554 pages; WJLA coverage) documented officials reclassifying violent assaults, robberies, and thefts under pressure from department leadership. The report issues findings on 18 officials, sustaining misconduct allegations against most of them, and states that all involved members were in a full duty status as of its March 5, 2026 date (p. 553); the subsequent discipline figures widely reported (19 officials charged, 13 facing termination) come from press coverage, not from the report itself. Federal prosecutors declined criminal charges (p. 20). The chief resigned in December 2025. The Washington Post's coverage frames the OIG report as the process audit explaining how that was possible. So this page asks two separately scoped questions: did the 2013 loss of oversight leave a mark, and did the documented 2023–2025 manipulation move the published totals.

This data sits inside a live dispute between federal and local officials over DC crime. This analysis does not adjudicate that dispute; it reports what the published data can and cannot show, with the confounds labeled.

Question 1: Did the 2013 loss of independent review shift the mix?

The question: When MPD's classification controls broke down in 2013, did the mix of offense categories shift?

The answer: Not attributably. 6 of 12 series do shift at the 2013 boundary after multiple-testing correction, but the identical test at placebo dates where nothing control-related happened fires comparably (up to 7 of 12). The classification mix shifts often throughout the archive, so the 2013 boundary cannot be singled out as a review-unit fingerprint. No evidence of a distinctive 2013 break is the result, and it is published as exactly that.

In July 2026 the DC Inspector General found that MPD disbanded its Staff Review Unit, the independent function that reviewed crime reports before they became official statistics, around 2013, and never replaced it (OIG No. 26-E-03-FA0, July 28, 2026). This page is different in kind from the eight disparity analyses: it audits the reliability of the incident archive itself. Our window extends earlier than the OIG's review period (2015–2025) because detecting a 2013 break requires data from before it. 6 of 12 series show a level shift at the January 2013 boundary that survives Newey-West errors and multiple-testing correction: theft (other) (+8.8 pp, 95% CI +2.1 pp to +15.5 pp); theft from auto (-8.4 pp, 95% CI -14.7 pp to -2.2 pp); assault w/ dangerous weapon (+0.6 pp, 95% CI +0.1 pp to +1.1 pp); theft from auto, as a share of all theft (-16.6 pp, 95% CI -28.7 pp to -4.6 pp); robbery, as a share of robbery + theft (-5.7 pp, 95% CI -9.4 pp to -2.0 pp); ADW, as a share of all violent categories (+2.8 pp, 95% CI +0.5 pp to +5.2 pp). But specificity fails: the identical test at placebo boundaries fires just as readily (up to 7 of 12 at dates where nothing control-related happened), so the honest reading is that the classification mix is generally nonstationary and the 2013 boundary does not stand out from that churn. This dataset provides no evidence of a distinctive review-unit fingerprint, and that null is the finding on the 2013 oversight question; the documented 2023–2025 manipulation era is tested separately below. This is a test of the published citywide totals, not an exoneration: case-level misclassification, and downgrades that move an incident out of the index categories entirely, would not appear here. The internal affairs report documents exactly that kind of case-level manipulation for 2023–2025, which is why Question 2 below tests that era directly.

6 of 12

Series (nine offense-category shares and three substitution pairs) with a level shift at January 2013 that survives Newey-West standard errors and Benjamini-Hochberg correction across the whole family. Read it with the placebo check below: a count here only means something if placebo dates fire less.

What this page cannot see. The public incident feed contains only the nine DC index-crime categories. A downgrade from an index offense to a non-index offense (for example, assault with a dangerous weapon recorded as simple assault) removes the incident from this dataset entirely. So this analysis can detect shifts BETWEEN visible categories and unexplained level changes IN a category, but it cannot directly observe exits from the feed. The internal affairs report puts primary numbers on exactly this invisible path: at least 406 documented theft reports were moved out of the feed (pp. 16, 217, 335, 370, 533), about 14 a month across the pressure era, and no aggregate test on this dataset can see them.
COVID-19Federal surgeReview unit disbanded (2013)Mark43 RMS goes live (Aug 2015)Pressure era begins (Aug 2023)Scandal breaks (Dec 2025)Classification training (Feb 2026)Theft (other)share of all monthly incidents · latest 50.3%2013 shift +8.8 pp (survives correction)Theft from autoshare of all monthly incidents · latest 22.7%2013 shift -8.4 pp (survives correction)Motor vehicle theftshare of all monthly incidents · latest 11.3%Robberyshare of all monthly incidents · latest 5.7%Burglaryshare of all monthly incidents · latest 4.8%Assault w/ dangerous weaponshare of all monthly incidents · latest 4.4%2013 shift +0.6 pp (survives correction)Sex abuseshare of all monthly incidents · latest 0.2%Homicideshare of all monthly incidents · latest 0.5%Arsonshare of all monthly incidents · latest 0.0%Theft from auto, as a share of all theftwithin-pair share · latest 31.1%2013 shift -16.6 pp (survives correction)Robbery, as a share of robbery + theftwithin-pair share · latest 10.2%2013 shift -5.7 pp (survives correction)ADW, as a share of all violent categorieswithin-pair share · latest 40.9%2013 shift +2.8 pp (survives correction)2008201020122014201620182020202220242026

Each row: the category's share of that month's incidents (blue) and, for the nine categories, the raw monthly count (gray, each on its own scale). The counts are shown because the nine shares sum to 1: a real surge in one category mechanically depresses every other share, so a share break can be an echo of a real change elsewhere. Yellow ticks mark detected breaks. The five dashed reference lines and the two shaded bands:

Post-reform data is thin. Only 7 complete months have accrued since the February 2026 training launch; nothing about the effect of the reforms can be concluded yet. The line is marked so the question is on the record, not because it is answerable.

The a-priori test: a level shift at January 2013

One fixed break point, chosen before looking, at the OIG-documented loss of independent review. Each series gets a segmented interrupted time-series regression (level shift AND slope change at January 2013, calendar-month dummies), estimated on a symmetric window of five years each side of the break (2008–2017), with Newey-West standard errors (4 lags) because monthly shares are autocorrelated and naive OLS intervals would be too narrow. The symmetric window is load-bearing twice over: it stops long-run trend curvature from loading onto the break term, and it ends before the 2020 COVID discontinuity and the 2023 motor vehicle theft wave, so neither can masquerade as a 2013 effect.

SeriesLevel shift at Jan 201395% CI (Newey-West)p (NW)p (BH-adjusted)Survives correction
Theft (other)+8.8 pp+2.1 pp to +15.5 pp0.0100.031Yes
Theft from auto-8.4 pp-14.7 pp to -2.2 pp0.0080.031Yes
Motor vehicle theft+0.9 pp+0.0 pp to +1.8 pp0.0450.078No
Robbery-0.7 pp-1.7 pp to +0.2 pp0.1300.195No
Burglary-1.2 pp-2.9 pp to +0.5 pp0.1610.214No
Assault w/ dangerous weapon+0.6 pp+0.1 pp to +1.1 pp0.0160.032Yes
Sex abuse+0.1 pp-0.1 pp to +0.3 pp0.2530.287No
Homicide+0.0 pp-0.0 pp to +0.1 pp0.2630.287No
Arson+0.0 pp-0.0 pp to +0.0 pp0.5150.515No
Theft from auto, as a share of all theft-16.6 pp-28.7 pp to -4.6 pp0.0070.031Yes
Robbery, as a share of robbery + theft-5.7 pp-9.4 pp to -2.0 pp0.0020.028Yes
ADW, as a share of all violent categories+2.8 pp+0.5 pp to +5.2 pp0.0160.032Yes

This is a family of 12 tests. At the 5% level, 12 uncorrected tests would be expected to produce about 0.6 false positives by chance alone, so no single uncorrected p-value below is headlined; the last column applies Benjamini-Hochberg across the family (q = 0.05). The three pair series use within-pair shares (for example, theft from auto as a share of all theft), which the compositional coupling between the nine global shares cannot move.

The specificity check: placebo break dates

The identical family test (same window shape, same errors, same correction) run at January boundaries where nothing control-related happened, all chosen so their windows stay clear of the COVID band. If placebo dates fire as often as 2013, the mix is churning generally and a 2013 count proves nothing about the review unit.

Break date testedCorrected discoveries
Jan 2013 (review unit disbanded)6 of 12
Jan 2011 (placebo)1 of 12
Jan 2012 (placebo)7 of 12
Jan 2014 (placebo)4 of 12
Jan 2015 (placebo)1 of 12

Placebo dates fire as often as the real boundary: the 2013 count carries no specificity.

Detected breaks, wherever they fall

Binary segmentation on each deseasonalized series, with no prior about where a break should be. A break landing on the Mark43 migration, in the COVID band, or in the federal-surge era rather than at a tested boundary is reported as exactly that.

SeriesDetected break month(s)
Theft (other)Jun 2009; Feb 2012 (near: Review unit disbanded (2013)); Feb 2014; Feb 2017; Mar 2019 (near: COVID-19 band begins (Mar 2020)); Apr 2020 (COVID-era; not attributable to classification practice); Nov 2023 (near: Pressure era begins (Aug 2023)); Aug 2025 (federal-surge era; nothing past this boundary is attributed to classification practice)
Theft from autoJul 2009; Feb 2011; Nov 2013 (near: Review unit disbanded (2013)); Feb 2016 (near: Mark43 RMS goes live (Aug 2015)); May 2020 (COVID-era; not attributable to classification practice); Nov 2022 (near: Pressure era begins (Aug 2023)); Sep 2025 (federal-surge era; nothing past this boundary is attributed to classification practice)
Motor vehicle theftJan 2009; Feb 2010; Feb 2011; Feb 2012 (near: Review unit disbanded (2013)); Jul 2020 (COVID-era; not attributable to classification practice); Nov 2022 (near: Pressure era begins (Aug 2023)); Nov 2023 (near: Pressure era begins (Aug 2023)); Jul 2025 (at the federal surge boundary; not attributable to classification practice)
RobberyJan 2014 (near: Review unit disbanded (2013)); Oct 2016; Jun 2020 (COVID-era; not attributable to classification practice); Mar 2023 (near: Pressure era begins (Aug 2023)); Oct 2024 (near: Federal surge begins (Aug 2025))
BurglaryAug 2010; Nov 2011; Dec 2013 (near: Review unit disbanded (2013)); Jan 2015 (near: Mark43 RMS goes live (Aug 2015)); Oct 2016; Apr 2022
Assault w/ dangerous weaponMay 2010; Sep 2016; May 2018; Apr 2020 (COVID-era; not attributable to classification practice); Dec 2021; Dec 2022 (near: Pressure era begins (Aug 2023)); Sep 2025 (federal-surge era; nothing past this boundary is attributed to classification practice)
Sex abuseDec 2011; Aug 2018; Jul 2024 (near: Pressure era begins (Aug 2023))
HomicideDec 2010; Mar 2018; Apr 2020 (COVID-era; not attributable to classification practice); Mar 2025 (near: Federal surge begins (Aug 2025))
ArsonApr 2009; Jul 2013 (near: Review unit disbanded (2013)); Mar 2016 (near: Mark43 RMS goes live (Aug 2015))
Theft from auto, as a share of all theftJun 2009; Feb 2017; Aug 2022 (near: Pressure era begins (Aug 2023)); Aug 2023 (near: Pressure era begins (Aug 2023)); Aug 2025 (federal-surge era; nothing past this boundary is attributed to classification practice)
Robbery, as a share of robbery + theftAug 2009; Mar 2012 (near: Review unit disbanded (2013)); Feb 2014; Jun 2016 (near: Mark43 RMS goes live (Aug 2015)); Apr 2020 (COVID-era; not attributable to classification practice); Feb 2023 (near: Pressure era begins (Aug 2023)); Oct 2024 (near: Federal surge begins (Aug 2025))
ADW, as a share of all violent categoriesMar 2010; Jan 2014 (near: Review unit disbanded (2013)); Jan 2022; Feb 2023 (near: Pressure era begins (Aug 2023)); Apr 2024 (near: Pressure era begins (Aug 2023)); Jul 2025 (at the federal surge boundary; not attributable to classification practice)

Question 2: Did the documented 2023–2025 manipulation move the published totals?

The question: During the era where manipulation is documented (2023–2025), did the published offense mix move the way the documented downgrade paths predict?

The answer: No aggregate signature. None of the four directional contrasts predicted by the documented downgrade paths survives correction. A null here does not mean manipulation did not happen: manipulation is documented at case level in the 554-page internal affairs report. It means the documented case-level manipulation was not large enough to visibly move citywide monthly aggregates in the public index-crime feed. Given the documented scale, a null is the expected result; the magnitude bridge below shows why.

Unlike 2013, this era comes with documented direction: the internal affairs report describes robberies logged as theft, weapons assaults downgraded below ADW, and thefts shifted to lesser categories. None of the four one-sided contrasts built from the internal affairs report's documented downgrade paths (robberies logged as theft, weapons assaults downgraded, thefts shifted to lesser categories) survives Benjamini-Hochberg correction at the August 2023 boundary. A null here does not mean manipulation did not happen: manipulation is documented at case level in the 554-page internal affairs report. It means the documented case-level manipulation was not large enough to visibly move citywide monthly aggregates in the public index-crime feed. The magnitude bridge below shows the documented scale falls below this test's detection floor, so this null is the expected result given the documented scale. The decline-then-reversion signature that aggregate-level manipulation would leave is ABSENT: no contrast both falls during the pressure era and reverts after exposure.

The era test reuses Question 1's machinery at the August 2023 boundary (the first full month of the tenure the House and internal affairs reports date the pressure to): segmented regression on a symmetric window of 24 months each side (Aug 2021–Jul 2025), which by construction avoids the COVID band and ends exactly at the August 2025 federal surge boundary. Newey-West errors throughout. Three families, each Benjamini-Hochberg corrected separately and labeled: the 12-series two-sided family, the 4-contrast one-sided directional family, and the 4-contrast reversion family.

The directional family (one-sided, signs fixed by the documented paths)

ContrastDocumented path (predicts a fall)Shift at Aug 2023p (one-sided, NW)p (BH)Fires
ADW, as a share of all violent categoriesweapons assaults downgraded below ADW (many exit the feed)-1.2 pp0.3810.550No
Robbery, as a share of robbery + theftrobberies logged as theft-0.7 pp0.4130.550No
Theft from auto, as a share of all theftthefts shifted toward the lesser theft category-2.1 pp0.0630.250No
Assault w/ dangerous weaponADW share of the whole mix falls as downgrades exit+0.2 pp0.8190.819No

The reversion test (post-exposure step, Dec 2025)

If the declines were manipulation, exposure should end them: the December 2025 House and DOJ reports and the resignation are the exposure boundary. Two caveats travel with this table: only 9 post-exposure months exist, and every one of them sits past the August 2025 federal surge boundary, so reversion cannot be cleanly attributed to exposure versus the surge. A post-exposure step with no matching pressure-era decline is NOT the signature: both legs are required, and a lone step can as easily reflect the surge-era mix.

ContrastPost-exposure step (Dec 2025)p (one-sided, NW)p (BH)Fires
ADW, as a share of all violent categories+2.8 pp0.1200.231No
Robbery, as a share of robbery + theft+1.6 pp0.1740.231No
Theft from auto, as a share of all theft+0.3 pp0.4420.442No
Assault w/ dangerous weapon+1.5 pp0.0010.005Yes

The magnitude bridge: could this test even see it?

The primary document (Final Investigative Report IS #26-000051, March 5, 2026) states no comprehensive total and calls its fullest tally "by no means a complete record" (p. 526), so the total here is our computed sum of the report's own per-official audit counts: Cmdr M. Pulliam 157 (p. 16); Capt Donigian 363 (p. 318); Capt Haskis 194 (p. 334); Capt Rivers 131 (p. 217); Capt R. Pulliam 47 (p. 370); Capt Merrill 30 (p. 533). By that count, the internal affairs investigation documents 922 improperly edited or reclassified reports across six named officials, the large majority of them theft reports moved out of the nine public index categories; the documented robbery and ADW reclassifications number 5 and 11 respectively. Press coverage's named figures (Pulliam 157, Donigian about 360) match the primary audits. For scale: the pre-surge era window averaged roughly 185 robberies, 87 ADW incidents, and 1070 thefts reported citywide per month.

Scenario Implied effect Minimum detectable effect (one-sided 5%)
Documented robbery path: 5 reports (pp. 16, 256, 309) 0.014 pp 5.25 pp on robbery share of robbery+theft
Documented ADW path: 11 reports (p. 16) 0.093 pp 6.44 pp on ADW share of violent
Documented theft exits: 406 reports counted (pp. 16, 217, 335, 370, 533), most of Donigian's 363 on top (p. 318) 14/month leave the feed (0.6% of monthly volume) Invisible by construction: exits never reach this dataset
Stress ceiling: all 922 edits as robbery downgrades 2.62 pp 5.25 pp
Stress ceiling: all 922 edits as ADW downgrades 8.72 pp 6.44 pp

Documented-scale manipulation falls below what this test can detect, so a null is the expected result given the documented scale. Even the stress ceiling, all 922 edits concentrated on the robbery path, stays under that contrast's detection floor; only if nearly all of them had been ADW downgrades would the shift have been visible, and the report's documented ADW count is 11.

The era boundary and its placebos (two-sided, all 12 series)

SeriesLevel shift at Aug 202395% CI (Newey-West)p (NW)p (BH-adjusted)Survives correction
Theft (other)+3.8 pp+0.4 pp to +7.1 pp0.0270.109No
Theft from auto+0.2 pp-3.0 pp to +3.4 pp0.8970.897No
Motor vehicle theft-5.1 pp-8.5 pp to -1.7 pp0.0030.042Yes
Robbery+0.4 pp-2.8 pp to +3.7 pp0.8010.897No
Burglary+0.2 pp-0.5 pp to +0.9 pp0.6150.897No
Assault w/ dangerous weapon+0.2 pp-0.3 pp to +0.7 pp0.3620.725No
Sex abuse+0.1 pp-0.2 pp to +0.5 pp0.5220.894No
Homicide+0.2 pp+0.0 pp to +0.3 pp0.0260.109No
Arson-0.0 pp-0.0 pp to +0.0 pp0.0770.231No
Theft from auto, as a share of all theft-2.1 pp-4.8 pp to +0.6 pp0.1250.300No
Robbery, as a share of robbery + theft-0.7 pp-7.0 pp to +5.5 pp0.8250.897No
ADW, as a share of all violent categories-1.2 pp-8.9 pp to +6.5 pp0.7620.897No
Break date testedCorrected discoveries
Aug 2023 (documented pressure era begins)1 of 12
Aug 2011 (placebo)1 of 12
Aug 2013 (placebo)0 of 12
Aug 2015 (placebo)1 of 12
Aug 2017 (placebo)0 of 12

Placebo dates fire as often as the real boundary: the Aug 2023 count carries no specificity.

Placebo limitation, stated: a window of this shape that avoids both the COVID band and the federal surge fits only the real boundary, so these placebos come from the pre-COVID era. They test the machinery's false-fire rate in calmer data, not a perfectly matched counterfactual.

The district-level test, and the trap it documents

The obvious next step after a citywide null is to de-dilute: about 21 documented edits a month vanish against thousands of citywide incidents, but concentrated into single districts the same edits loom much larger. The internal affairs report names the districts and the windows itself, and one cell looked large enough to see from outside: the Seventh District, January to October 2024, where the report's audit window and per-official counts put roughly 15 theft reports a month leaving the public feed against a base of about 75 (pp. 88, 217, 546). We built that test, pre-registered its interpretation before estimating (database/theft_exit.py, wording committed before the estimator existed), and ran it.

The design does not hold, and that is the finding. A difference-in-differences estimate is only evidence if the treated district moved in parallel with the comparison districts before the window; the Seventh District did not. Its theft counts rose 38% in 2023 (806 to 1,113), the steepest rise in the city that year, and then fell 16% in 2024. The comparison districts are not a quiet backdrop either: over the same year they ranged from -0.3% to +31%, so the pre-window period is one where districts were moving sharply and differently from one another. Weighing the whole twelve-month lead path against the same placebo assignments the estimate would use, the pre-registered permutation joint lead test rejects parallel pre-trends at p = 0.016 against a gate of 0.10. The one cell whose documented scale approached visibility therefore cannot be estimated, and the remaining cells sit below their detection floors. The pre-registered branch for this outcome (design-fails) says it plainly: the honest output is that answer, not a weaker substitute test.

The trap, said plainly, because someone else will run this. A difference-in-differences on public data without a pre-trend check reports a large, statistically significant 2024 theft deficit in the Seventh District, in precisely the district and window the internal affairs report documents. It would look like independent corroboration of the documented manipulation. It would be an artifact: the district reverting from its own 2023 spike, measured against a baseline year it was never tracking. The chart below is that evidence. The deficit a naive test finds is the downslope of a hill that was already there before the window opened.
Audit window (Jan–Oct 2024, p. 217)0 = moving in step with comparison districtsFirst 7D THRIVE area (May 2024)2023: already above the comparison, and rising, before the window opens2023-012023-072024-012024-072025-012025-07

Seventh District theft (theft/other plus theft from auto), monthly, as log-point deviations from its own late-2022 baseline relative to the comparison districts, with a 95% band. The audit window (p. 217) is shaded; the first THRIVE area in the district (May 2024, p. 13) is the dashed line, marked separately so the window-versus-initiative timing stays readable. The divergence is the 2023 lead-up, before either boundary.

Two further failures converge on the same verdict, one statistical and one supplied by the report itself. First, every documented cell sits below its placebo-derived detection floor, the Seventh District marginally so:

Documented cellDocumented edits/moBase thefts/moExpected dipDetection floorVerdict
Seventh District, January-October 202415.17516.7%18.4%below detection floor
Second District, January-August 202515.43584.1%14.4%below detection floor
Fifth District (Rosedale THRIVE), February-August 20257.42323.1%15.3%below detection floor

Detection floors here come from the placebo distribution itself (the deficit size placebo district-windows produce 5% of the time), which at seven districts is wider than asymptotic approximations suggest. A cell below its floor cannot confirm or rule out the documented pattern; its null is uninformative by construction.

Second, the report's own findings place commander-directed Seventh District misclassification across 2023 and 2024 (pp. 88, 541–542), which means the comparison baseline year is itself a documented editing year: the design has no clean pre-period available in that district for reasons the primary document supplies. The pre-trend verdicts:

CellPermutation joint lead pPlacebo assignmentsGate (fails at p ≤ 0.10)
Seventh District, January-October 20240.01663FAIL
Second District, January-August 20250.73718pass
Fifth District (Rosedale THRIVE), February-August 20250.42118pass

Composed with the citywide result above, this is the substantive finding: at both resolutions available in public data, citywide aggregates and district-level contrasts, documented case-level manipulation leaves no testable trace. The manipulation is established by the internal affairs report's sustained findings, not by this page; what this page establishes is that the public feed could not have caught it, and that the one analysis that looks like it catches it is a trap. With no measurable deficit in the Seventh District cell, the window-versus-THRIVE timing comparison has nothing to separate, and no timing claim is made.

How this number is built (and where it's soft)

Related: Enforcement vs reported crime  |  Methodology  |  All analyses