How 232 campgrounds were counted twice
· 3 min · CampsiteFit

For two days in August this site listed Live Oak South twice.
Two entries, identical coordinates, ninety sites each, the same ninety sites. Search returned both, and every count that touched California counted them both. So did a hundred and eighty-four other campgrounds.
What happened
The federal reservation API is queried one state at a time. Arizona's listing returns 499 rows for 459 distinct facilities, because the API's own paging hands some records back more than once.
The importer read its map of existing campgrounds from disk once, before the
run started. So when a repeat arrived, it looked for a campground it had already
written during that same run, did not find one on disk, decided the name was
taken, and filed the second copy as live-oak-south-2.
185 facility identifiers ended up under two names each. 232 of them were published. 3,610 campsites were counted twice.
What it broke
Everything downstream, quietly:
- Search showed the same campground twice, with identical data.
- State counts ran high, California, Arizona, Texas and Alaska worst.
- Two published articles quoted the inflated figures.
Nothing failed, and no test caught it. The pages rendered correctly, the numbers added up internally, and the corpus was self-consistently wrong.
What found it was an odd-looking pair of rows in a table. They shared coordinates, and then they turned out to share a facility identifier, which is the thing that should have made them one record.
The fix, and why it is boring
The importer now registers each campground as it writes it, so a repeat inside the same run resolves to the record just created and compares as unchanged rather than as new.
A cleanup pass removed the copies already on disk. It kept whichever copy held more campsites rather than assuming the two matched, because dropping data to fix a duplicate would have been the worse bug, and it reported any group that disagreed. None of the 185 did.
Since then, 1,624 new fact files across 26 states have produced zero duplicates. That is the fix tested at four times the scale of the failure. A gate now reloads the corpus and fingerprints the result, so a second load that produces different rows from the first fails the build instead of shipping.
Why publish this
Because the site's argument is that a published number is worth checking, and that obliges us to say when ours were wrong.
Both affected articles were recomputed and carry a dated note saying what changed. One of them, where a 38 ft fifth wheel actually fits, had its whole state table rebuilt. The figures gate that now guards them registers every published number with the query that produces it and checks the two on every build. It exists because of an earlier incident of exactly this kind, where an import silently falsified two live articles and nothing failed.
There is a pattern in the defects this project keeps finding, and this is one of them: the data was correct and the code did not read it. The facility identifier sat in every record from the first import, and nothing compared it. The access field went the same way, for months.
The counter-measure is not care. It is asking, on a schedule, what the source publishes that nothing here reads, a question that has now produced four corrections and will produce more.
The 499 rows, the 459 facilities and the 185 identifiers all come from the Recreation Information Database, which is public domain, and the corrected records are the ones every campground page links to today.