Census & Kinship Math

How the NARA Soundex Phonetic Code Works for Surname Variants

Family names do not always stay the same on paper. See how Soundex groups similar sounds and where you still need to search separately.

Illustrative scene of handwritten family records, a magnifying glass, and a paper family tree.
Illustrative archive scene.

At a glance

  • American Soundex converts any surname into a 4-character alphanumeric code (LDDD) consisting of the surname's initial letter followed by three digits representing six articulatory consonant classes.
  • Under the official National Archives (NARA) rule, adjacent consonants with the same code collapse into a single digit even if separated by H or W (Ashcraft → A261), whereas vowels (A, E, I, O, U, Y) act as separators that force the consonant digit to be coded twice (Tymczak → T522).
  • Because Soundex locks in the initial letter of the surname, phonetically identical names with different first letters (such as Cline C450 vs. Kline K450, or prefix names like VanDeusen V532 vs. Deusen D250) require parallel code searches.

Before the 20th century, American surname spelling was fluid, phonetic, and dependent on the ear of the clerk or census enumerator holding the pen. The same family might appear as Smith, Smyth, Smythe, or Schmid across four consecutive decades, or as Meyer, Meier, Maier, and Myer on neighboring county schedules. If historical records were indexed strictly alphabetically, a single swapped vowel or doubled consonant would hide an ancestor hundreds of pages away.

To solve this problem, Robert C. Russell and Margaret King Odell patented the American Soundex phonetic indexing algorithm (US Patents 1,261,167 in 1918 and 1,435,663 in 1922). Beginning in 1935 under the Works Progress Administration (WPA), the federal government used Soundex to index household cards across the 1880, 1900, 1910, 1920, and 1930 United States Federal Censuses so older citizens could prove their age under the newly enacted Social Security Act. Every variant of Smith, Smyth, Smythe, and Schmid collapses into a single four-character code: S530. You can generate verified NARA Soundex codes—with a letter-by-letter trace—in Tab C of our Tombstone, Census & Kinship Calculator.


How the WPA and NARA Indexed the 1880–1930 US Censuses

As the National Archives (NARA) Soundex guide documents, WPA clerks transcribed census schedules onto 3 × 5 inch index cards organized first by state, then by 4-character Soundex surname code (LDDD), and finally alphabetically by the household head’s given name. Each decennial census has distinct Soundex coverage rules that affect your search strategy:

Census Year NARA Index Type Geographic & Household Coverage Key Indexing Nuance
1880 Soundex (T734–T780) All states/territories, ONLY households with a child aged 10 or under Households without a resident child aged ≤ 10 were omitted from the 1880 WPA Soundex cards; check Enumeration District schedules directly for childless households.
1900 Soundex (T1030–T1083) 100% of households nationwide across all states and territories Complete every-household index; includes birth month and birth year (June 1, 1900 statutory date).
1910 Soundex or Miracode 21 states indexed (AL, AR, CA, FL, GA, IL, KS, KY, LA, MI, MO, MS, NC, OH, OK, PA, SC, TN, TX, VA, WV) Miracode uses the exact same LDDD Soundex code for the surname, paired with microfilm visitation numbers instead of page/line numbers.
1920 Soundex (M1548–M1605) 100% of households nationwide across all states and territories Complete every-household index (January 1, 1920 statutory date).
1930 Soundex (M2050–M2061) 12 Southern states (AL, AR, FL, GA, LA, MS, NC, SC, TN, VA, plus 7 counties in KY and 7 in WV) WPA funding ended before Northern, Midwestern, and Western states were card-indexed.

The 6-Class Articulatory Consonant Table

Every NARA Soundex code has the fixed format LDDD: one uppercase letter (A–Z) followed by three digits (0–6) representing six articulatory consonant groups:

Soundex Digit Consonant Letters Articulatory Phonetics Group Common Interchange Examples
1 B, F, P, V Bilabial & Labiodental stops/fricatives (lips and front teeth) Pfeiffer / Fifer, Bauer / Vauer
2 C, G, J, K, Q, S, X, Z Sibilants, Affricates & Velar/Palatal stops (hissing and hard/soft K/G sounds) Cook / Koch, Schwartz / Swartz, Jackson / Jaxon
3 D, T Alveolar/Dental stops (tongue tip behind upper teeth) Millerd / Millert, Schmidt / Schmid
4 L Lateral alveolar liquid (voiced L sound) Kelly / Kelley
5 M, N Bilabial & Alveolar nasals (frequently confused in cursive minims) Samson / Sanson, Bruns / Brums
6 R Rhotic liquid (voiced R sound) Barker / Barger
Separators A, E, I, O, U, Y Vowels & Semivowel Y — dropped after letter 1, but separate identical consonant classes Tymczak (C/Z and K separated by A → both coded as 2)
Bridges H, W Glottal & Labiovelar Approximants — ignored after letter 1; bridge identical consonant classes Ashcraft (S and C bridged across H → single 2)

The Exact NARA Adjacent-Duplicate Rule: Why H and W Bridge While Vowels Separate

The most common bug in simplified Soundex calculators (including standard SQL SOUNDEX() functions that follow simplified textbook rules instead of the strict NARA specification) is mishandling adjacent consonants that share the same Soundex digit (1–6). Here is the exact five-step NARA algorithm:

  1. Retain the First Letter (L) and Note Its Numeric Class: Strip any punctuation or apostrophes (O'Brien → OBrien), write down the first letter of the surname in uppercase (L), and look up its Soundex class (1–6, or 0 for vowels/H/W) so you can test whether the second letter is an immediate duplicate.
  2. Collapse Immediate Adjacent Duplicates (Including Against the First Letter): Two adjacent consonants from the same Soundex class (LL, SS, CK, SC, DT, MN, PF) are coded as one single digit. If the second letter of the surname has the same class as the first letter (Pfister: P = 1, f = 1; Lloyd: L = 4, l = 4; Gutierrez: G and t/r/z), the second letter is dropped without emitting a digit.
  3. The H and W Bridge Rule (Ashcraft → A261): Under the official NARA specification, if two consonants with the same Soundex digit are separated only by H or W, the second consonant is NOT coded. You can think of H and W as completely transparent: strip H and W before testing for adjacent duplicate consonant codes. In Ashcraft, S (2) and C (2) are separated only by H (S-H-C), so they merge into a single 2, producing A261 (whereas a non-NARA implementation that fails to bridge across H erroneously outputs A226).
  4. The Vowel Separator Rule (Tymczak → T522): If two consonants with the same Soundex digit are separated by a vowel (A, E, I, O, U, Y), the second consonant IS coded. Vowels are dropped from the final code, but they reset the duplicate guard. In Tymczak, c (2) and z (2) merge into one 2, but the vowel a sits between cz (2) and k (2), forcing k to be coded as a second 2 (T522).
  5. Truncate or Zero-Pad to Three Digits (LDDD): Keep only the first three consonant digits (Washington → W252), or pad short names on the right with zeros (Lee → L000, Jackson → J250).

Five Step-by-Step Worked NARA Soundex Encodings

1. Washington → W252 (Standard Truncation & H Ignore)

Letters:   W    a    s    h    i    n    g    t    o    n
Classes:  [0]  (v)   2   (h)  (v)   5    2    3   (v)   5
Action:    W   drop  2   ign  drop  5    2    3   drop  5  --> Raw: W25235 --> Truncate: W252

2. Pfister → P236 (Initial-Letter Adjacent Duplicate Collapse)

Letters:   P    f    i    s    t    e    r
Classes:  [1]   1   (v)   2    3   (v)   6
Action:    P   dup! drop  2    3   drop  6  --> Exact 3 digits: P236

Because initial P (1) and second letter f (1) share class 1 with no vowel between them, f is dropped.

3. Ashcraft → A261 (H Bridges Identical Consonant Codes)

Letters:   A    s    h    c    r    a    f    t
Classes:  [0]   2   (h)   2    6   (v)   1    3
Action:    A    2   ign  dup!  6   drop  1    3  --> Raw: A2613 --> Truncate: A261

Because h is transparent, c (2) merges into s (2), allowing r (6) and f (1) to fill the code (A261).

4. Tymczak → T522 (Adjacent Merge cz + Vowel Separator a Before k)

Letters:   T    y    m    c    z    a    k
Classes:  [3]  (v)   5    2    2   (v)   2
Action:    T   drop  5    2   dup! sep   2  --> Exact 3 digits: T522

Here y is dropped, m = 5, c and z merge into the first 2, and the vowel a separates cz from k, coding k as the second 2 (T522).

5. Jackson → J250 (Vowel Separates Initial J from Triple-Cluster cks)

Letters:   J    a    c    k    s    o    n
Classes:  [2]  (v)   2    2    2   (v)   5
Action:    J   sep   2   dup! dup! drop  5  --> Raw: J25 --> Zero-pad: J250

Although J, c, k, s all belong to class 2, the vowel a separates initial J from c, so cks is coded as 2, followed by n = 5 and trailing 0 (J250).


Prefix Surnames (Van/Von/De/Di/La/Le and Mac/Mc) and Initial-Letter Variants

  1. Prefix Surnames (Van, Von, De, Di, La, Le, O'): NARA instructed clerks to index prefix surnames both with and without the prefix because different enumerators wrote them as one word or two words. Always search both codes: for VanDeusen, check V532 (with prefix) and D250 (Deusen, without prefix); for DeLaCruz, check D426, L262, and C620; for LeBlanc, check L145 and B452. Scottish and Irish Mac and Mc surnames (MacDonald / McDonald) both encode to M235, though some WPA microfilm rolls filed Mc- cards in a separate alphabetical sub-drawer at the front of the M sequence before Ma-.
  2. Initial-Letter Substitutions: Because Soundex never alters the first letter of the surname, any phonetic shift on letter 1 moves the card to a different letter drawer even when the three trailing digits stay identical. Always run parallel searches for common initial pairs: Cline (C450) vs. Kline (K450), Schwartz (S632) vs. Zwartz (Z632), Fifer (F160) vs. Pfeiffer (P160), and silent initial letters like Knudsen (K532) vs. Nudsen (N325) or Wright (W623) vs. Rite (R300).

Once you locate your ancestor’s card, narrow their birth date with US Federal Census Dates (1790–1950) and Birth Year Math, verify DNA matches using Cousin Chart and Shared DNA Centimorgan (cM) Math Explained, transcribe faded schedules with Using AI to Transcribe Old Cursive Wills, Deeds, and Letters, and scan fragile paper records using our Photo, Slide & Home Movie Digitization Planner.

Put it into practice

Try it with your own collection

Sources & further reading

  1. The Soundex Indexing System - National Archives (NARA)
  2. US Federal Census Records Overview - National Archives (NARA)
  3. Soundex - Wikipedia

Information on this page is for educational archival preservation and historical genealogy research. Always test conservation handling on non-unique materials first, verify AI handwriting transcriptions against original county or NARA microfilm, and never use autosomal DNA statistics for clinical or legal parentage determinations. Nothing on this site is legal, probate, medical, or financial advice. Spotted an error? Tell us and we will review it under our corrections policy.

Back to the beginning ↑