One name search downloads as a single Word file of 500 to 3,000 articles. Much of it is the same story, republished by eight different newspapers. Much of the rest is a different human being who happens to share the name. Clearing that out by hand is a staffer’s week with a highlighter. We built software that does it instead — 29 client searches, 33,771 articles, a written reason attached to every removal.
Every figure here comes from a real client search in the 2026 cycle. Candidates are described by office and state rather than named, and no number has been rounded in our favor.
What comes back is not a research file. It is a search result, and a search result has no idea what you were looking for. The junk comes in two shapes that fail in opposite directions: one buries the researcher in things they have already read, the other quietly hands them facts about a stranger.
Across the fourteen searches where the duplicate pass was logged article by article, 7,457 of 23,792 were copies of another article sitting in the same file. In one Iowa House race it was 46% — nearly half of what came back was the file repeating itself.
In one Florida House race, 231 of the 481 articles the search returned were about a different human being. Not similar people — honor‑roll students, high‑school track athletes, two student newspaper reporters, and the lead singer of an Irish rock band.
One story from a national news service gets republished three to eight times, each time under a different newspaper’s name. A newspaper chain runs a single piece across dozens of its own papers over several weeks. What that came to on fourteen real searches:
| Race | Articles returned | Duplicates removed | Share |
|---|---|---|---|
| Iowa House seat | 638 | 293 | 46% |
| Arizona governor | 1,999 | 844 | 42% |
| Florida Senate | 2,000 | 840 | 42% |
| Colorado House seat | 2,997 | 1,176 | 39% |
| Arizona House seat | 2,071 | 776 | 37% |
| Pennsylvania House seat | 2,051 | 756 | 37% |
| Second Iowa House seat | 2,049 | 718 | 35% |
| Second Pennsylvania House seat | 2,500 | 784 | 31% |
| Second Colorado House seat | 1,244 | 369 | 30% |
| Ohio House seat | 1,813 | 538 | 30% |
| Nevada governor | 840 | 222 | 26% |
| Third Iowa House seat | 112 | 24 | 21% |
| Second Florida House seat | 481 | 56 | 12% |
| Florida House seat | 2,997 | 61 | 2% |
| These fourteen searches | 23,792 | 7,457 | 31% |
Between a quarter and a half of every search is the same material twice. The three searches at the bottom are the exception that explains the rule: in each of them most of the file had already been removed as somebody else, so there was very little left to duplicate. In the second Iowa race, 245 original stories accounted for 985 articles — four copies apiece on average, and the most-copied ran far higher.
The candidate’s surname is also an ordinary English word for a job, and it is shared by several accomplished people. A search cannot tell the difference. Of 2,997 articles: 611 were about a different real person, 766 used the surname as an occupation, and 1,434 mentioned the name without saying anything about him. Another 71 were export cover sheets, foreign-language reprints, or copies of an article already in the file.
Every one of the other 2,882 articles is still there, filed with a note naming who it was actually about. The client can read that list. Nothing was thrown away.
Duplicates run at a steady third of every search. Articles about the wrong person do not: the share depends entirely on how common the candidate’s name is, and there is almost nothing in the middle. Eleven of those searches carry a categorized exclusion log that names, for every article removed, the person it turned out to be about:
| Race | Articles returned | Not this person | Share |
|---|---|---|---|
| Florida House seat | 2,997 | 2,821 | 94% |
| Third Iowa House seat | 112 | 68 | 61% |
| Second Florida House seat | 481 | 271 | 56% |
| Arizona governor | 1,999 | 83 | 4% |
| Arizona House seat | 2,071 | 85 | 4% |
| Pennsylvania House seat | 2,051 | 57 | 3% |
| Nevada governor | 840 | 21 | 3% |
| Colorado House seat | 2,997 | 41 | 1% |
| Second Colorado House seat | 1,244 | 18 | 1% |
| Second Pennsylvania House seat | 2,500 | 35 | 1% |
| Florida Senate | 2,000 | 27 | 1% |
| These eleven searches | 19,292 | 3,527 | 18% |
Put both removals together and 9,435 of those 19,292 articles came out — 49%. 5,908 were copies of an article already in the file; 3,527 were about somebody else, or never said anything about the subject at all. What is left, 9,857 articles, is the file the researcher actually reads.
We run this on client searches today, and we are bringing it into the Civly platform so any Lexis subscriber can hand over a raw file and get back a clean, citable set of articles. Nothing is deleted — every removed article is kept and listed with the reason it went, so you can always say what is and is not in the public record. If you have a search sitting on a shared drive that nobody has had time to read, that is the one we want to see.
Measured on 29 client LexisNexis searches processed during the 2026 cycle, 33,771 articles parsed in total · the 31% duplicate figure covers the fourteen searches whose duplicate pass was logged article by article (7,457 of 23,792) · the 18% wrong-person share and the 49% headline both cover the eleven searches that carry a categorized exclusion log naming the person behind every removal (3,527 and 9,435 of 19,292 respectively) · candidates are identified by office and state only, and every figure is drawn from the engagement’s own removal log rather than an estimate.