News Search Cleanup · Research Brief
← Newsroom
News Search Cleanup · What a name search actually returns

49% of what a Lexis search returns is a duplicate, or somebody else entirely.

One name search downloads as a single Word file of 500 to 3,000 articles. Much of it is the same story, republished by eight different newspapers. Much of the rest is a different human being who happens to share the name. Clearing that out by hand is a staffer’s week with a highlighter. We built software that does it instead — 29 client searches, 33,771 articles, a written reason attached to every removal.

By · Chief Technology Officer, Civly ·
Client searches cleaned
29
Articles reviewed
33,771
Duplicate or wrong person
49%
Reached the research file
51%

Every figure here comes from a real client search in the 2026 cycle. Candidates are described by office and state rather than named, and no number has been rounded in our favor.

The cost nobody quotes you

Two kinds of junk, and someone has to clear both by hand

What comes back is not a research file. It is a search result, and a search result has no idea what you were looking for. The junk comes in two shapes that fail in opposite directions: one buries the researcher in things they have already read, the other quietly hands them facts about a stranger.

The same story, again
31% are duplicates

The same story, in three to eight newspapers

Across the fourteen searches where the duplicate pass was logged article by article, 7,457 of 23,792 were copies of another article sitting in the same file. In one Iowa House race it was 46% — nearly half of what came back was the file repeating itself.

Somebody else entirely
231 of 481 articles

The search matched the letters, not the person

In one Florida House race, 231 of the 481 articles the search returned were about a different human being. Not similar people — honor‑roll students, high‑school track athletes, two student newspaper reporters, and the lead singer of an Irish rock band.

How much of it is the same story

About a third of every search is the same story, over and over

One story from a national news service gets republished three to eight times, each time under a different newspaper’s name. A newspaper chain runs a single piece across dozens of its own papers over several weeks. What that came to on fourteen real searches:

RaceArticles returnedDuplicates removedShare
Iowa House seat63829346%
Arizona governor1,99984442%
Florida Senate2,00084042%
Colorado House seat2,9971,17639%
Arizona House seat2,07177637%
Pennsylvania House seat2,05175637%
Second Iowa House seat2,04971835%
Second Pennsylvania House seat2,50078431%
Second Colorado House seat1,24436930%
Ohio House seat1,81353830%
Nevada governor84022226%
Third Iowa House seat1122421%
Second Florida House seat4815612%
Florida House seat2,997612%
These fourteen searches23,7927,45731%

Between a quarter and a half of every search is the same material twice. The three searches at the bottom are the exception that explains the rule: in each of them most of the file had already been removed as somebody else, so there was very little left to duplicate. In the second Iowa race, 245 original stories accounted for 985 articles — four copies apiece on average, and the most-copied ran far higher.

When the name isn't the person

One Florida House race returned 2,997 articles. 115 made the research file.

The candidate’s surname is also an ordinary English word for a job, and it is shared by several accomplished people. A search cannot tell the difference. Of 2,997 articles: 611 were about a different real person, 766 used the surname as an occupation, and 1,434 mentioned the name without saying anything about him. Another 71 were export cover sheets, foreign-language reprints, or copies of an article already in the file.

Under 4% of a paid search was the person who paid for it

2,997articles returned 115reached the research file

Every one of the other 2,882 articles is still there, filed with a note naming who it was actually about. The client can read that list. Nothing was thrown away.

How much is somebody else

Articles about somebody else can be 1% of a search, or 94% of it

Duplicates run at a steady third of every search. Articles about the wrong person do not: the share depends entirely on how common the candidate’s name is, and there is almost nothing in the middle. Eleven of those searches carry a categorized exclusion log that names, for every article removed, the person it turned out to be about:

RaceArticles returnedNot this personShare
Florida House seat2,9972,82194%
Third Iowa House seat1126861%
Second Florida House seat48127156%
Arizona governor1,999834%
Arizona House seat2,071854%
Pennsylvania House seat2,051573%
Nevada governor840213%
Colorado House seat2,997411%
Second Colorado House seat1,244181%
Second Pennsylvania House seat2,500351%
Florida Senate2,000271%
These eleven searches19,2923,52718%

Put both removals together and 9,435 of those 19,292 articles came out — 49%. 5,908 were copies of an article already in the file; 3,527 were about somebody else, or never said anything about the subject at all. What is left, 9,857 articles, is the file the researcher actually reads.

Put it to work

You already paid for the articles. Get the research file.

We run this on client searches today, and we are bringing it into the Civly platform so any Lexis subscriber can hand over a raw file and get back a clean, citable set of articles. Nothing is deleted — every removed article is kept and listed with the reason it went, so you can always say what is and is not in the public record. If you have a search sitting on a shared drive that nobody has had time to read, that is the one we want to see.


Measured on 29 client LexisNexis searches processed during the 2026 cycle, 33,771 articles parsed in total · the 31% duplicate figure covers the fourteen searches whose duplicate pass was logged article by article (7,457 of 23,792) · the 18% wrong-person share and the 49% headline both cover the eleven searches that carry a categorized exclusion log naming the person behind every removal (3,527 and 9,435 of 19,292 respectively) · candidates are identified by office and state only, and every figure is drawn from the engagement’s own removal log rather than an estimate.