Census & Kinship Math
How the NARA Soundex Phonetic Code Works for Surname Variants
Family names do not always stay the same on paper. See how Soundex groups similar sounds and where you still need to search separately.

At a glance
- American Soundex converts any surname into a 4-character alphanumeric code (LDDD) consisting of the surname's initial letter followed by three digits representing six articulatory consonant classes.
- Under the official National Archives (NARA) rule, adjacent consonants with the same code collapse into a single digit even if separated by H or W (Ashcraft → A261), whereas vowels (A, E, I, O, U, Y) act as separators that force the consonant digit to be coded twice (Tymczak → T522).
- Because Soundex locks in the initial letter of the surname, phonetically identical names with different first letters (such as Cline C450 vs. Kline K450, or prefix names like VanDeusen V532 vs. Deusen D250) require parallel code searches.
Before the 20th century, American surname spelling was fluid, phonetic, and dependent on the ear of the clerk or census enumerator holding the pen. The same family might appear as Smith, Smyth, Smythe, or Schmid across four consecutive decades, or as Meyer, Meier, Maier, and Myer on neighboring county schedules. If historical records were indexed strictly alphabetically, a single swapped vowel or doubled consonant would hide an ancestor hundreds of pages away.
To solve this problem, Robert C. Russell and Margaret King Odell patented the American Soundex phonetic indexing algorithm (US Patents 1,261,167 in 1918 and 1,435,663 in 1922). Beginning in 1935 under the Works Progress Administration (WPA), the federal government used Soundex to index household cards across the 1880, 1900, 1910, 1920, and 1930 United States Federal Censuses so older citizens could prove their age under the newly enacted Social Security Act. Every variant of Smith, Smyth, Smythe, and Schmid collapses into a single four-character code: S530. You can generate verified NARA Soundex codes—with a letter-by-letter trace—in Tab C of our Tombstone, Census & Kinship Calculator.
How the WPA and NARA Indexed the 1880–1930 US Censuses
As the National Archives (NARA) Soundex guide documents, WPA clerks transcribed census schedules onto 3 × 5 inch index cards organized first by state, then by 4-character Soundex surname code (LDDD), and finally alphabetically by the household head’s given name. Each decennial census has distinct Soundex coverage rules that affect your search strategy:
| Census Year | NARA Index Type | Geographic & Household Coverage | Key Indexing Nuance |
|---|---|---|---|
| 1880 | Soundex (T734–T780) |
All states/territories, ONLY households with a child aged 10 or under | Households without a resident child aged ≤ 10 were omitted from the 1880 WPA Soundex cards; check Enumeration District schedules directly for childless households. |
| 1900 | Soundex (T1030–T1083) |
100% of households nationwide across all states and territories | Complete every-household index; includes birth month and birth year (June 1, 1900 statutory date). |
| 1910 | Soundex or Miracode | 21 states indexed (AL, AR, CA, FL, GA, IL, KS, KY, LA, MI, MO, MS, NC, OH, OK, PA, SC, TN, TX, VA, WV) |
Miracode uses the exact same LDDD Soundex code for the surname, paired with microfilm visitation numbers instead of page/line numbers. |
| 1920 | Soundex (M1548–M1605) |
100% of households nationwide across all states and territories | Complete every-household index (January 1, 1920 statutory date). |
| 1930 | Soundex (M2050–M2061) |
12 Southern states (AL, AR, FL, GA, LA, MS, NC, SC, TN, VA, plus 7 counties in KY and 7 in WV) |
WPA funding ended before Northern, Midwestern, and Western states were card-indexed. |
The 6-Class Articulatory Consonant Table
Every NARA Soundex code has the fixed format LDDD: one uppercase letter (A–Z) followed by three digits (0–6) representing six articulatory consonant groups:
| Soundex Digit | Consonant Letters | Articulatory Phonetics Group | Common Interchange Examples |
|---|---|---|---|
1 |
B, F, P, V |
Bilabial & Labiodental stops/fricatives (lips and front teeth) | Pfeiffer / Fifer, Bauer / Vauer |
2 |
C, G, J, K, Q, S, X, Z |
Sibilants, Affricates & Velar/Palatal stops (hissing and hard/soft K/G sounds) |
Cook / Koch, Schwartz / Swartz, Jackson / Jaxon |
3 |
D, T |
Alveolar/Dental stops (tongue tip behind upper teeth) | Millerd / Millert, Schmidt / Schmid |
4 |
L |
Lateral alveolar liquid (voiced L sound) |
Kelly / Kelley |
5 |
M, N |
Bilabial & Alveolar nasals (frequently confused in cursive minims) | Samson / Sanson, Bruns / Brums |
6 |
R |
Rhotic liquid (voiced R sound) |
Barker / Barger |
| Separators | A, E, I, O, U, Y |
Vowels & Semivowel Y — dropped after letter 1, but separate identical consonant classes |
Tymczak (C/Z and K separated by A → both coded as 2) |
| Bridges | H, W |
Glottal & Labiovelar Approximants — ignored after letter 1; bridge identical consonant classes | Ashcraft (S and C bridged across H → single 2) |
The Exact NARA Adjacent-Duplicate Rule: Why H and W Bridge While Vowels Separate
The most common bug in simplified Soundex calculators (including standard SQL SOUNDEX() functions that follow simplified textbook rules instead of the strict NARA specification) is mishandling adjacent consonants that share the same Soundex digit (1–6). Here is the exact five-step NARA algorithm:
- Retain the First Letter (
L) and Note Its Numeric Class: Strip any punctuation or apostrophes (O'Brien→OBrien), write down the first letter of the surname in uppercase (L), and look up its Soundex class (1–6, or0for vowels/H/W) so you can test whether the second letter is an immediate duplicate. - Collapse Immediate Adjacent Duplicates (Including Against the First Letter): Two adjacent consonants from the same Soundex class (
LL,SS,CK,SC,DT,MN,PF) are coded as one single digit. If the second letter of the surname has the same class as the first letter (Pfister:P = 1, f = 1;Lloyd:L = 4, l = 4;Gutierrez:Gandt/r/z), the second letter is dropped without emitting a digit. - The
HandWBridge Rule (Ashcraft → A261): Under the official NARA specification, if two consonants with the same Soundex digit are separated only byHorW, the second consonant is NOT coded. You can think ofHandWas completely transparent: stripHandWbefore testing for adjacent duplicate consonant codes. InAshcraft,S(2) andC(2) are separated only byH(S-H-C), so they merge into a single2, producingA261(whereas a non-NARA implementation that fails to bridge acrossHerroneously outputsA226). - The Vowel Separator Rule (
Tymczak → T522): If two consonants with the same Soundex digit are separated by a vowel (A, E, I, O, U, Y), the second consonant IS coded. Vowels are dropped from the final code, but they reset the duplicate guard. InTymczak,c(2) andz(2) merge into one2, but the vowelasits betweencz(2) andk(2), forcingkto be coded as a second2(T522). - Truncate or Zero-Pad to Three Digits (
LDDD): Keep only the first three consonant digits (Washington → W252), or pad short names on the right with zeros (Lee → L000,Jackson → J250).
Five Step-by-Step Worked NARA Soundex Encodings
1. Washington → W252 (Standard Truncation & H Ignore)
Letters: W a s h i n g t o n
Classes: [0] (v) 2 (h) (v) 5 2 3 (v) 5
Action: W drop 2 ign drop 5 2 3 drop 5 --> Raw: W25235 --> Truncate: W252
2. Pfister → P236 (Initial-Letter Adjacent Duplicate Collapse)
Letters: P f i s t e r
Classes: [1] 1 (v) 2 3 (v) 6
Action: P dup! drop 2 3 drop 6 --> Exact 3 digits: P236
Because initial P (1) and second letter f (1) share class 1 with no vowel between them, f is dropped.
3. Ashcraft → A261 (H Bridges Identical Consonant Codes)
Letters: A s h c r a f t
Classes: [0] 2 (h) 2 6 (v) 1 3
Action: A 2 ign dup! 6 drop 1 3 --> Raw: A2613 --> Truncate: A261
Because h is transparent, c (2) merges into s (2), allowing r (6) and f (1) to fill the code (A261).
4. Tymczak → T522 (Adjacent Merge cz + Vowel Separator a Before k)
Letters: T y m c z a k
Classes: [3] (v) 5 2 2 (v) 2
Action: T drop 5 2 dup! sep 2 --> Exact 3 digits: T522
Here y is dropped, m = 5, c and z merge into the first 2, and the vowel a separates cz from k, coding k as the second 2 (T522).
5. Jackson → J250 (Vowel Separates Initial J from Triple-Cluster cks)
Letters: J a c k s o n
Classes: [2] (v) 2 2 2 (v) 5
Action: J sep 2 dup! dup! drop 5 --> Raw: J25 --> Zero-pad: J250
Although J, c, k, s all belong to class 2, the vowel a separates initial J from c, so cks is coded as 2, followed by n = 5 and trailing 0 (J250).
Prefix Surnames (Van/Von/De/Di/La/Le and Mac/Mc) and Initial-Letter Variants
- Prefix Surnames (
Van,Von,De,Di,La,Le,O'): NARA instructed clerks to index prefix surnames both with and without the prefix because different enumerators wrote them as one word or two words. Always search both codes: forVanDeusen, checkV532(with prefix) andD250(Deusen, without prefix); forDeLaCruz, checkD426,L262, andC620; forLeBlanc, checkL145andB452. Scottish and IrishMacandMcsurnames (MacDonald / McDonald) both encode toM235, though some WPA microfilm rolls filedMc-cards in a separate alphabetical sub-drawer at the front of theMsequence beforeMa-. - Initial-Letter Substitutions: Because Soundex never alters the first letter of the surname, any phonetic shift on letter 1 moves the card to a different letter drawer even when the three trailing digits stay identical. Always run parallel searches for common initial pairs:
Cline(C450) vs.Kline(K450),Schwartz(S632) vs.Zwartz(Z632),Fifer(F160) vs.Pfeiffer(P160), and silent initial letters likeKnudsen(K532) vs.Nudsen(N325) orWright(W623) vs.Rite(R300).
Once you locate your ancestor’s card, narrow their birth date with US Federal Census Dates (1790–1950) and Birth Year Math, verify DNA matches using Cousin Chart and Shared DNA Centimorgan (cM) Math Explained, transcribe faded schedules with Using AI to Transcribe Old Cursive Wills, Deeds, and Letters, and scan fragile paper records using our Photo, Slide & Home Movie Digitization Planner.
Put it into practice
Try it with your own collection
Family history calculator
Compare dates, explore census clues, and calculate family relationships.
Sources & further reading
Information on this page is for educational archival preservation and historical genealogy research. Always test conservation handling on non-unique materials first, verify AI handwriting transcriptions against original county or NARA microfilm, and never use autosomal DNA statistics for clinical or legal parentage determinations. Nothing on this site is legal, probate, medical, or financial advice. Spotted an error? Tell us and we will review it under our corrections policy.
Back to the beginning ↑