UC Berkeley researchers have assembled local laws from all 50 states into a free, open-access database. The project uses AI and data science to make millions of ordinances easier to search and compare.
The team published a version online in June, and some users have already built searchable interfaces. Last week, a paper describing the work was accepted for the Conference on Neural Information Processing Systems.
Local Laws Become Machine-Readable Data
According to the project report, the archive covers every digitally accessible law collected by the researchers. It focuses specifically on municipal and county rules rather than state or federal statutes.

Lead researcher Diag Davenport is an assistant professor of technology policy, governance and society at Berkeley. Davenport began considering the project while studying algorithms, systemic bias and the criminal legal system.
The research question required comparing local laws in places with different histories of discrimination. Davenport found no database containing all those laws.
Instead, documents were spread across proprietary databases and inconsistent government websites. The scale included more than 3,000 counties and another 6,000 or so jurisdictions making their own laws.

“Once you realize how fragmented it all is, it’s easy to understand why no one’s done the work,” Davenport said.
AI Processes Roughly 7 Million Pages
Davenport worked with Denis Peskoff, a Berkeley postdoctoral scholar, on collecting the documents. They consulted lawyers and designed the process around the hosting websites’ technical requirements.
AI researcher Joe Barrow and Berkeley undergraduate Christopher Vu also contributed. The larger technical challenge was converting nearly 10,000 documents into usable data.
The documents contained roughly 7 million pages, including blurry or poorly structured PDFs. The team used LightOnOCR, a vision-language optical character recognition model, to extract text and document structure.
The resulting machine-readable text is easier and less costly to search and analyze than the original PDFs. Researchers then used OpenAI models to tag samples and organize the local laws into categories across the corpus.
The database contains millions of unique ordinances. It lets researchers identify patterns in rules governing housing, public space, business activity and everyday conduct.
Related technology coverage examines complications created by digital tools. Block Editorial has also reported on AI access to company systems.
Public Tools Could Compare Local Laws
Davenport said developers could use the repository to train chatbots for specific subjects. A tool focused on Bay Area building codes could compare apartment construction requirements in Berkeley and El Cerrito.
That could help developers estimate project costs or homeowners understand backyard accessory dwelling unit rules. Those uses remain examples of possible tools built from the data.
For broader housing coverage, readers can also see costs of moving between states.
Davenport said forthcoming tools based on the dataset must be free and publicly accessible. Journalists, researchers and policymakers could use it to compare regional rules and examine how understandable laws are.
The paper, titled “Freeing the Law with LOCUS: A Local Ordinance Corpus for the United States,” is available on arXiv. The researchers will present their findings in December.





