About the data
Thank you for visiting Durham in Data! I have been a resident since 2014 and am raising my two children here. Both attend DCS. I work from home and have a marketing business that has been around for more than 20 years. The world is changing really fast and everyone is busier than ever trying to keep up. I created this site leveraging the power of modern AI to bring vast volumes of public data together and try to make some sense of it. No one has time to watch board meetings. Maybe some are able to get to the odd town meeting or glance at a commission's minutes on the town website. I don't have the time or interest in these things but I do have a great deal of interest in understanding what is happening around me in this community.
Sources have been meticulously documented and cross-referenced. You'll find source links everywhere. I employed modern data visualization techniques and reviewed most of the pages to ensure they made sense.
This was a labor of love, and I hope you find it as useful and insightful as I did in creating it. We might not all agree on decisions our elected officials make, but if we can all get the same factual information, we're best equipped to form informed opinions. That is what this site is about.
This site is a structured index of public information about the Town of Durham, Maine, drawn from the town's own records and from state and federal sources. It is not affiliated with the town. It carries no commentary and argues for no outcome: a figure, the source it came from, and the method used to get there.
Everything here was already public.
This site is the culmination of hours of time, using AI and compiling and surfacing information in public sources.
What this is
Everything here was already public. The work is in gathering it into one place, checking each figure against the source it came from, and labelling what it does and does not show. 740 town documents (6,403 pages and 1.89 million words, of which 93 are scans that had to be read by OCR) sit behind the pages on this site, alongside 11 datasets from 5 publishing bodies. The sources page lists every one of them: why it is included, how much of it this site holds, and where to get it yourself.
The town publishes the records; this site indexes them. Where this site and the town's own copy of a document disagree, the town's copy is the record and the figure here is the error.
The site is built and maintained independently and in a private capacity, and holds no official standing of any kind.
What this site does not do
It does not publish property-level tax data. The town's annual commitment book lists the name and assessed value attached to every property in Durham. It is public information, and this project holds it: the Taxes page reports aggregates drawn from it, including the largest owners and the most valuable parcels, which is ordinary civic reporting. The full per-account table is deliberately not published. Turning it into a searchable index of who owns what and what they pay is a decision a person should make on purpose, not a side effect of a pipeline.
It does not take a position. Where two figures diverge (the school funding formula charging towns by property valuation rather than pupil count, for instance), the divergence is stated as the sources state it. That is arithmetic, not an accusation.
It does not put words in anyone's mouth. The meeting transcripts come from automatic speech recognition, which records that the speaker changed but never who was speaking. Every name attached to a statement is therefore checked against the attendance list in the town's written minutes; where no record confirms it, the name is published with a (?) after it. 186 of 349 meetings could be matched to minutes naming who attended. No video is rehosted; each item links to the second it came from on the town's own channel.
How figures are checked
Checks run at build time
Every figure names its source and its method
That is what the Source & method toggle under each chart contains: the body that published the numbers, and how they were derived from what it published: which column was read, which rows were dropped, what was recomputed. A figure that cannot name both does not go on the site.
Numbers scraped from documents are candidates, not findings
Budget presentations state hypothetical rates in exactly the same words as real ones ("using this year's assessed values, the mil rate would be"), so a regular expression that finds a number in a PDF has found a string, not a fact. Scraped figures are held as candidates with their surrounding context for a person to confirm. What gets published is either confirmed against minutes or balanced against the source's own internal arithmetic: a school funding year appears only when six identities inside the ED279 report reconcile to the cent, including the municipal allocations summing to the district total.
Survey estimates carry their margin of error
One source here is a survey rather than a record. The American Community Survey estimates a town of about 4,262 residents from a sample of its households; it does not count them. Every estimate is published with the margin the Census Bureau publishes beside it, and a change is called real only when it is larger than the margins on both ends of the comparison. Of the 28 comparisons tested on the People page, 9 clear that test. The other 19 are reported as no measurable change rather than written up as trends.
Two counts of the corpus, deliberately different
The document corpus is counted twice, and the two counts do not match on purpose. 782 PDF files were text-extracted from the town's website and its published Drive folders, 96 of them through OCR because they are scans. That is a count of files on disk. Every figure on this site uses the other count: 740 distinct documents, which is what remains after 29 byte-identical duplicates are dropped (93 of those documents are image-only). 782 files minus 29 duplicates is 740 documents.
Removing the duplicates is not housekeeping. 25 of the 29 removed copies come from one folder, the 2016 Select Board minutes, where 60 files are only 35 distinct documents, the rest being second copies of files already there under a slightly different name. Counting files rather than documents would have shown that year holding nearly twice the meetings it actually held, and the spike would have looked like a finding.
Source & method
durhammaine.gov and the town's YouTube channel
782 PDF files text-extracted, 96 of them via OCR (counts of files on disk, before deduplication). The coverage figures count documents: 740 distinct PDFs remain once 29 byte-identical duplicates are removed, 93 of them image-only. Videos enumerated, not downloaded
https://durhammaine.gov/documents
Data as of 1 September 2026 (the newer of the documents and meetings dates); retrieved 7 September 2026.
Corrections
Any figure here can be checked without taking this site's word for it, and that is the intended remedy for an error. Every chart carries a Source & method toggle naming the publisher and the derivation. The sources page links each publisher's own copy of the data, so the original can be opened alongside what is shown here. The search box covers every page here, including the extracted text of 335 town documents and all 349 meeting transcripts. That text is machine-read, so a document quoted on this site has to be confirmed against the town's own copy, which every document page links. The 410 historical minutes and budgets from the town's Drive folders are not published as pages and are outside the search.
Where this site disagrees with the body that published the data, the publisher is right. A figure that cannot be reproduced from its stated source is an error on this site, and should be treated as one.
An error that survives that check is worth reporting. Write to inquiry@durhamindata.org, quoting the page and the figure. A correction that can be traced to a source will be made; where the source itself is wrong, that will be said on the page rather than quietly adjusted.
