Wednesday, August 19, 2026

Logic & Fallacies in Genealogy - An AI Perspective


Last week I wrote a blog post titled Logic & Fallacies in Genealogy. That post was an offshoot of a presentation that I have been working on for several months highlighting the potential errors in logic that genealogists may encounter during their research. At the end of that article I promised a follow-up article on how AI introduces additional errors into your research and how to deal with them. As I have been using AI more frequently in my research I have noticed that the AI platforms frequently come to conclusions without the supporting data. I asked the AI platform to identify the errors in logic that have occurred as it produces conclusions. Sometimes the AI platforms have caught their errors but the majority of the errors were caught by me. I am working with the AI to catalog the errors and develop ways to resolve or reduce their frequency. The results have been interesting.

Let's start at the beginning. I have found it helpful to upload various background documents prior to starting a research project in AI. These documents set the stage for what is to come. The background documents may include a locality guide to provide background on the resources available for the area, the role that AI will take in the process (research assistant) and its expertise, and a list of logical fallacies to be aware of (see the previous post for the list). I then ask AI to produce a research plan, research log, and a tracking sheet of the errors that occur during the research project. Each of these documents is updated regularly throughout the process. 

I will be focusing on the error tracking sheet in this blog post. The following are the result of three individual projects spanning several weeks of research in multiple sessions.

Foundations: What the Error Tracking Sheet Established

1) Reasoning from Unverified Readings

This was the most common and costly error across the projects. This error consisted of the AI building structural claims on readings that were marked uncertain and never tested. Examples include misread digits (reading a year, age, or date incorrectly), misinterpreted initials or letters, and outlier readings treated as significant (a last name misspelled or incorrectly listed in one record). 

The fix: Ask AI to list all information and compare it in a table. Ask AI to place uncertain readings in brackets [...]. This gives you the opportunity to decide on the correct interpretation.

2) Jurisdiction Assumed Rather Than Established

This error happened frequently during these research projects. The AI repeatedly inferred jurisdiction from the nearest named place rather than the subject's documented location. For example, parish records indicated that the individuals attended St. Augustine Catholic Church in Minster, Ohio. The AI inferred that this meant the individuals lived in Minster, Auglaize County and the AI created research plan focused on Auglaize County records. However, the individuals lived across the county boarder in Shelby County. Additionally, Auglaize County was established in 1848 from Mercer County. The AI did not consider that early parishioners would have records in Mercer County.

The fix: Township, county, parish, and birthplace are distinct fields that must be independently verified. Clearly identify the timeline of locality formation, records locations, and potential for living in different jurisdictions.

3) Predicting Document Content from General Practice

AI repeatedly assumed that records would contain specific information such as age, heirs, parents, and place of birth, based on previous records. For example, the AI assumed that marriage records would have the parents' names or that naturalization records would have the person's place of birth. These predictions failed multiple times based on local record keeping practices. This demonstrates the need to confirm record structure before applying methods.

The fix: Do not allow the AI to assume record content without reviewing the specific records. Provide a baseline rule that records need to be reviewed prior to assuming content.

4) Argument from Ignorance

Nulls or negative results were repeatedly treated as evidence about the past rather than evidence about the finding aid or record. Every negative finding has three possible explanations: the event did not happen; it happened but was not recorded; it was recorded but the finding aid cannot reach it. There were several cases where this occurred in these projects. One was during a cholera epidemic. The parish death register did not list the several hundred people who died of cholera during 1849-1850 with the burials so AI assumed that these deaths were not recorded. On the contrary, there was a separate list which just listed the names and month of death for these victims.

The fix: Make sure the AI is aware of the types of records available for an area. Are the church records more or less complete than the civil records? Are there overlapping volumes where the same dates are recorded in more than one place? Are there actual gaps in the record keeping?

5) Unexamined Premises About Sources

Assumptions about what a record series must contain (or does not contain) were repeatedly overturned by sampling adjacent pages. A property observed on one page is a property of that page only and is not the rule for all other pages. This error may occur when different people are providing individual pages, i.e., a new clerk or priest versus the previous clerk or priest or a different enumerator in the census. 

The fix: Let AI know that the handwriting has changed or the format of the entries have changed as you review pages in the records. Different formats may include more or less information than previous record formats.

6) Pseudo-replication

Multiple observations depending on a single contested reading were counted as independent witnesses. One example that occurred frequently was establishing the death certificate, obituary, and headstone as individual pieces of evidence when there is a high likelihood that the informant on the death certificate also provided the information for the obituary and the headstone.

The fix: Establish a rule to count witnesses, not documents, and state certainty against the weakest link.

7) Failure of Execution

This is not a fallacy per se but it does cause research problems. This consists of naming a check but not running it and then using the ledger to outrank the evidence. The AI may suggest that a specific record be reviewed to check the accuracy of information from another document based on the assumption the record will have the information. It then takes the assumption that hasn't been proven and treats it as evidence.

The fix: Either ensure that all suggested records are reviewed or indicate that the specific record has not been reviewed and that no conclusions can be inferred from it until that check has been performed.

Additional Errors that May Occur

Conclusion by preponderance of trees

I will provide information from online family trees as a baseline document for a project. These may include family group sheets with sources or screen shots from FamilySearch, Ancestry, etc. However, I always provide a caveat with this information highlighting that they are online trees and subject to error. I also ask AI to review them and point out all conflicts, problems, and unsupported information prior to moving forward.

Name similarity as identity

AI may interpret people with the same or similar names as a single individual. As the researcher, you need to point out when there are several people with the same name in a locality and that they may appear in the same records. Make sure the AI is informed when information about an individual is being added as information versus as a specific detail for the research subject.

Age arithmetic as a fact

AI will do the math if you provide a record such as a death record that states the person was 76y 5m 13d to determine the birth date. I have had records where the birth record was a couple days, months, or even a year off from the math and AI inferred they were different individuals. Remember, you are the researcher and you make the final decision on what is the correct information.

Naming convention overreach

AI is aware of typical naming conventions and will use it to invent grandparents or other relatives. Be aware of this and tell AI that names will be based on the records researched and that a naming convention is a potential research clue, not a fact.

Confident sounding plausibility

AI will convince itself that something is a fact because it is plausible. I always provide AI with a rule that conclusions are based on the quality of the record. The AI is instructed to determine if the source is original, derivative, or authored; the information is primary, secondary, or undetermined; and the evidence is direct, indirect, or negative. The AI is also told to assess the quality of any conclusions it makes by classifying it as possible, plausible, probable, highly probable, proven, or disproven. 

Invented citations

AI can and will invent citations that fit the conclusion. Make sure that you look up every citation the AI presents and tell the AI when something does not exist. It is your responsibility to push back and not accept everything the AI gives you.

Plausible transcription of illegible text

AI does a good job of transcribing text but it does depend on the quality of the record to begin with. It will make up words to fill in the gaps if it is allowed. I once had an AI produce an entire paragraph in a Will that was not actually there. When having AI do transcriptions you should tell it to produce the transcription exactly as the document is written, maintain all punctuation, line breaks, and spelling errors, and indicate any uncertain text with brackets [...]. You can then help the AI by providing your reading of any bracketted text or indicating where it made errors.

Safeguards that Actually Work

Provide this list to AI at the start of each project and set them as rules to follow.
  • Label speculation immediately
  • State refutation conditions in advance
  • Record negative findings with reasons they may be false
  • Survey surrounding pages before accepting nulls
  • Control for informant reliability

Pre-Statement Checklist

Evaluate the AI responses with the following in mind:

  1. Is any part built on an uncertain reading? Is the uncertainty marked on the claim?
  2. Does the evidence require a list or comparison table rather than a selection?
  3. Is jurisdiction established from the subject’s location?
  4. Is the AI predicting record content? Do the fields exist in the record?
  5. Is the AI weighting a null? Has the finding-aid's/record's coverage been established?
  6. Is the AI generalizing the record content from one page or enumerator?
  7. How many independent witnesses are there really?
  8. Is the AI applying a group-level pattern to an individual?
  9. Is the AI stating a prediction without also applying a refutation condition?
  10. Is this fact, inference, or speculation? Has it been labelled as such?

Closing

Across projects, most errors were caught externally by the researcher, not the AI. The discipline’s strength lies not in preventing errors but in bounding their cost. The combined framework above provides a unified, method-focused reference for preventing propagation, designing tests, and maintaining evidential integrity.

I hope that some of this helps you as you explore the use of AI in your genealogy research and gives you the comfort to push back on or accept the input AI gives.

Wednesday, August 12, 2026

Logic & Fallacies in Genealogy

Logic & Fallacies in Genealogy - That sounds like a complex problem. You may be asking - "What do I know about logic and how does it effect my view of the genealogy research that I do?" Well, I have been thinking the same thing lately. How are the conclusions that I make driven by the records I find versus the preconceived notions I have? Am I understanding the full value of everything that a record tells me or am I inferring something that isn't actually there? And how would I know if I have done that?

Logic in genealogy empowers researchers to distinguish truth from tradition using structured thinking. Understanding logical validity (how conclusions follow from premises) vs truth (what reflects reality) is key. Deductive reasoning applies general principles to specific cases; inductive reasoning builds generalizations from data. Common fallacies, like confirmation bias or post hoc reasoning, can skew genealogical analysis. Clear logic helps us write responsible, credible reports and builds stronger ancestral narratives. Does this sound like something you vaguely remember from a college class but have forgotten over the years? Don't feel bad, most genealogists will fall into at least one of the traps I will outline below sometime during their research. Understanding the causes and how they impact your research is the first step in avoiding them.

Why does logic matter in genealogy?

  • It helps us avoid the temptation of wishful thinking in ancestral discovery.
  • Logical rigor strengthens the credibility of our research and improves our ability to tell the story.

Have you ever experienced a case where sloppy reasoning led to a faulty family connection? Maybe you took a family story and then forced a record to fit that narrative. Maybe you found a family where the individuals were similar to the family you were researching and forced the record to fit your expectations. Or, maybe you accepted a hint from your favorite website and added it without fully vetting the information. I believe that we have all made mistakes in our research at some time.

What determines truth versus falsehood in research?

  • The source of the information is important. Is the source original, derivative, or authored? Is the information from that source primary, secondary or undetermined?
  • How reliable is the record? The age, proximity, bias, or transcription of the record can all add or detract from the reliability of the record.
  • Is it truth or a consensus? Just because 30 online trees say the same thing doesn't make it true. 

Here is a real-life example of something I ran across this week. I was researching a man named John Fischer, born in Germany in 1826 who had immigrated to Ohio. The majority of family trees on Ancestry, MyHeritage, and FamilySearch have him as part of a family from Marienfeld, Westphalia. The family has 8 other children between 1815 and 1834. All of the children have exactly one source attached, the church baptismal records from FamilySearch, but no images of those records are available on FamilySearch. The only child without a baptism record attached was John Fischer. What evidence is there for this person to be a child in this family? Absolutely none! But he is carried over in dozens of family trees across multiple platforms.

Logical Validity versus Actual Truth

Can something be logically valid but not actually true? The simple answer is Yes. Logical validity is a conclusion that follows the premises. In your research you may find a census record that indicates a person was born in 1892. Upon further research you find his memorial on FindAGrave and it has 1892 carved in stone. You now have two sources that support the conclusion that he was born in 1892. But was he? How many additional sources do you need to find the truth? 

Let's look at my great-grandfather, Ray Westerheide. I have at least 16 sources for his birth. 

  • There is a county birth register that lists his birth as 12 December 1894.
    • This record is attached as a source twice because it is in DGS #005328962 and #004017413. Are these two sources or one? Obviously it is only one source but it has two citations for the source. 
    • How accurate is this record? It was produced near the time of the event. We hope it was provided by a responsible representative, maybe the parents or the doctor/midwife who delivered the baby. However, almost every entry on the page is written in the same hand which means these are not the original report. An error could have been made when transferring information to the county register.
  • His birth certificate states his birth date as 12 December 1894.
    • That is the same date as the county birth register. However, we need to examine this record closely. It is a certified copy of the birth certificate and is dated 18 September 1962. That means it was produced by referring to another record, likely the county birth register. 
    • This record also provides an additional clue. The date when the original record was recorded was 31 March 1895, over 3 months after the birth date on the record. This adds some room for doubt for this record because it was written after the event and transcribed from a record that was not an original.
  • His baptismal record states his birth was 4 December 1894 and he was baptized on 12 December 1894.
    • The baptism date on this record matches the birth date on the county birth register.
    • The birth date matches the date given on his headstone, death certificate, and obituary but are those actually three independent sources? They are likely all provided by the same source so can't be considered independent.
    • One additional thing to note. The baptism record was provided on 12 September 1962.
  • His WW I draft registration states his birth date as 4 December 1895. 
    • Who supplied this information? We would assume that he supplied it since he signed the document.
    • Why is the birth year 1895 when his birth record and baptism record say 1894?
  • The 1920 census gives his age as 25, which infers his birth about 1895.
  • The 1930 census gives an age of 34, which infers his birth about 1896.
  • The 1940 census says he is 43, which infers a birth around 1897. 
    • We know that his wife provided this information since she has the (x) next to her name.
  • He registered for the WW II draft in April 1942 and listed his age as 46 with a birth date of 4 December 1896. 
    • He signed this document so we assume that he provided the information. But why is it one year off from what he stated in the WW I draft card and two years off from his birth records?
  • The 1950 census lists his age as 54, which infers a birth about 1896.
  • FindAGrave and BillionGraves have his birth date as 4 December 1894 since that is the date carved on his headstone. 
    • This date is exactly one year earlier than the WW I draft registration, two years earlier than the WW II draft registration, and the same month and year as the county birth register.
  • His death certificate, obituary, and Social Security Death Index agree with the 4 December 1894 birth date.
    • The informant for the death certificate was Laura Westerheide, his 84 year old wife.

So, when was Ray born? Logically, depending on which record I had found, any of those dates would be valid. But truthfully, only one can be correct. Which one would you choose and why?

Deductive versus Inductive Reasoning in Genealogy

Deductive reasoning takes a general rule and produces a specific conclusion. An example of deductive reasoning is that naming conventions are found in many countries. The children are often named after their grandparents. Knowing this general rule, you might consider that the oldest son is named after one of the grandfathers and the oldest daughter is named after a grandmother. But is this always the case? Could it lead you down the wrong research path?

Inductive reasoning uses specific examples to produce a general conclusion. An example of inductive reasoning occurs when you are doing census research and find several families on adjacent pages with the same last name. You might conclude that these families are related. You haven't decided how they are related but it might be worth further research.

Other Common Logical Fallacies

Ad hoc reasoning is when you make up an explanation to suit the evidence. Have you ever decided that even though the record has conflicts it must be your family. The conflicts can easily be dismissed because the record keeper just made a mistake or maybe the neighbor provided the information. One example I find often is based on German church records. The child may be named Johann Heinrich in the baptism record but later you find a marriage record for Johann Frederick who would have been born the same year in the same town. That's close enough, right? It must be the same person so let's add him to the tree. I spend hours searching for the records, attaching them to individuals, and untangling families that were created on FamilySearch based on this type of assumption.

Another common fallacy is appealing to tradition. "That is how it has always been told." Have you ever let the family lore drive your research? Did you discover that it was not fully based on truth? My wife had a family story that one of her ancestors was English nobility. She left England because she was supposed to marry a Russian prince but she didn't want to marry him. She ran away from home and ended up marrying a French sea captain. That's a great story but the truth was a little less dramatic. The truth was that her parents died when she was a minor. She ended up living with her brother and his wife who had also taken in the wife's siblings. She married one of the wife's brothers. One of her sisters also married another of the wife's brothers. 

  • How did we get the nobility part of the story? We believe it was because her brother was elected to the House of Commons and served for several terms. 
  • How did we get the French sea captain? Her husband was a ship captain from the Channel Islands. They were a shipping family with a long history of maritime occupations.
  • How did we get the Russian prince? We have no idea.

Post hoc fallacy is caused by assuming cause just because it came before. One common example is frequently stated for Irish immigrants from the mid-19th century. The Irish Famine occurred between 1845 to 1852. Many Irish immigrants came to the United States during the 1850s. Some, but not all, came because of the famine. Attributing their immigration to the famine without supporting evidence is an example of post hoc fallacy. They may have come over as part of a chain migration to join family already in the US. They may have arrived as indentured laborers. They may have emigrated on their own volition. Or maybe we don't know the actual reason.

The last type of logical fallacy that I will present today is confirmation bias. Have you ever selectively chosen to weigh one record above another because it supported your theory? Cherry-picking evidence or skipping over records because they don't fit your expectations is an example of confirmation bias. You can avoid this by writing down alternative possibilities, not just your favored theories and see how the records support the theories.

To avoid these logical fallacies when doing your research you should:

  • state assumptions clearly,
  • acknowledge uncertainties,
  • avoid conclusive language when the evidence is incomplete,
  • use footnotes to explain reasoning,
  • and include alternative interpretations when applicable.

In conclusion, genealogy isn't just name-chasing - it is detective work with discipline. Knowing where and how logical fallacies can occur in your research is the first step in addressing them and improving your overall research success.

I hope to write at least one more post on logic and fallacies in genealogy in the near future. I am working with AI in my research. The AI frequently comes to conclusions without the supporting date. As these conclusions occur, I am cataloging them and developing ways to resolve or reduce their frequency. I hope to put this together as a blog post in the near future. Until then, I wish you luck in your research.