Gender API
Sign Up Free
Language expand_more

How to choose a gender API

Every provider in this category will tell you it is accurate, and none of them measured it on your list. Here is what actually separates them — and our own answer under each, so you can hold us to the same questions.

Last reviewed 2026-09-01 Written and maintained by the Gender-API.com team

The short answer

  • Compare on what one result hands you — a bare label cannot be thresholded; a probability plus a sample count can. That decides whether you automate part of the file or all of it by hand.
  • Ignore headline accuracy percentages, including ours — we do not publish one, because we have not measured it on a published benchmark. Measure it yourself on a few hundred of your own rows instead.
  • A bigger name database is not automatically better: a name counts the same whether three records back it or fifty thousand.
  • Then check the four things people forget — localization by country, stability of the answer over years, what comes back for an unknown name, and where the data is processed.

Start with what a single result gives you

This is the criterion that changes how much work you have to do, and it is visible in a provider's example response before you sign anything.

The same lookup returned three ways. A bare gender label gives you nothing to threshold on, so every row is either trusted blind or reviewed. Adding a probability lets you set a confidence cut-off, but not tell a 0.98 backed by three records from a 0.98 backed by three thousand. Adding a sample count as well tells you whether the confidence rests on enough evidence to believe. The useful question is not how accurate a provider is, but what a single result hands you to decide with.
The difference is not accuracy — it is whether you can act on one row without looking at all of them.

A gender label on its own is a verdict with no working shown. You can either apply it everywhere or check it everywhere; there is no middle. A probability gives you a cut-off. A probability with the number of records behind it tells you whether that cut-off means anything for the row in front of you — because a confident-looking answer resting on three records and one resting on thousands look identical until somebody shows you the count.

Our answer: every result carries a probability and a sample count, and the CSV/Excel output writes them into your file as columns, so the same threshold works whether you are calling the API or opening a spreadsheet.

Six dimensions worth checking

Six places providers in this category tend to differ, with what we do beside each. The left-hand column is what to look for — we have not audited anyone else, so none of it is a claim about a particular provider. Take it to whoever is on your shortlist, including us.

Six dimensions worth checking with any name-gender provider, each paired with what we do. One answer per name regardless of country, against localization by country, locale or IP — Andrea is male in Italy and female in Germany. A label with nothing behind it, against a probability and a sample count on every result. A large name count with no per-name evidence, against 9,124,598 names where each result carries its own count. Evaluation behind a sales call, against 100 free lookups a month with no card. Capacity that expires or an annual tie-in, against prepaid credits that never expire or a monthly plan cancellable at any time. An English-only interface, against a site and documentation in 16 languages.
The left column is what to check for, not a claim about anyone in particular — we have not audited other providers.

On the last two: credits bought as a one-off package do not expire, and a monthly plan renews monthly and can be cancelled at any time (switching package means cancelling and picking another). Both are on the pricing page in full, without a quote.

Why a coverage number does not settle it

Two names in the same database. Both count towards a headline coverage figure identically, but one rests on tens of thousands of records and can be acted on automatically, while the other rests on three and should be held back. A total name count cannot tell them apart; only the per-result sample count can.
Why a total name count is a poor thing to choose on.

We hold 9,124,598 names across 192 countries, and we put those figures on the homepage like everyone else does. They are worth exactly this much: they tell you how often you will get an answer at all, and nothing about whether a particular answer is worth using. Do not choose on them — ours or anyone's.

The test that does settle it costs an afternoon: take a few hundred rows from your own data where you already know the right answer, include the ones you expect to be hard, and run them through every provider on your shortlist. The names everybody gets right tell you nothing.

The checklist

Nine questions, in the order they tend to matter. Our answer is in the right-hand column — ask the same of anyone else you are looking at, and ask for the answers in writing rather than from a feature grid.

Ask Why it matters Our answer
What does one result contain? Decides whether you can automate the confident rows and review only the rest Gender, a probability, and the number of records behind it
Can the same name answer differently by country? Andrea is male in Italy and female in Germany; a single global answer is wrong in one of them Yes — by country code, browser locale or IP address
Will the same request answer the same way next year? An audit asks how a record was classified, and a moving answer cannot be explained Yes — it is arithmetic over stored records, not a generated guess
What comes back for a name it does not know? An empty field can be filtered; a plausible guess cannot be spotted An explicit "not found" and a sample count that says why — never an invented answer
Is there a route that needs no developer? The person with the list is usually not the person who can call an API CSV and Excel upload with the original workbook preserved
Where is the data processed, and is there a DPA? A list of customer names is personal data, and procurement will ask before you launch German company, servers in Germany, processing in the EU, DPA on request
Does unused capacity expire? A one-off clean-up should not need a subscription that outlives it Prepaid credits, and they do not expire
Can I test it before talking to anyone? If evaluating needs a sales call, you cannot compare shortlists in an afternoon 100 free lookups a month, no card, no call
What is the published accuracy? A single percentage describes the name set it was measured on, which is not yours We do not publish one — see below

The accuracy figure we will not give you

You will be quoted percentages in this market. We have no measured one to quote: there is no published benchmark of ours on a labelled name set, so any figure we printed would be a marketing number wearing a lab coat. We would rather say that than dress one up.

It is also less useful than it sounds. Accuracy on a name-gender lookup depends almost entirely on which names you ask about — a list of common German first names and a list of transliterated surnames from a dozen scripts will not produce the same number from any provider alive. So a single percentage tells you about the set it was measured on, and your list is not that set.

When you are quoted one, the question that makes it meaningful is: measured on which names, how many, labelled by whom, and counting "unknown" as what? If those answers are not available, the percentage is decoration.

How to run the evaluation in an afternoon

This page keeps telling you to measure it on your own data, so here is the method. It takes an afternoon, it works the same for every provider on your shortlist, and step five is the one that decides the project.

  • 1. Pull 200–300 rows you already know the answer for. From your real data, not a list of famous names. Deliberately include the awkward ones: rare names, non-Western names, surname-first entries, single-word names, hyphenated and married names.
  • 2. Do not clean them up first. The trailing whitespace, the titles, the "Dr." and the ALL CAPS are part of what you are testing. A provider that only works on tidy input has not solved your problem.
  • 3. Keep the whole response, not just the label. You will need the probability and the sample count in step five, and if a provider does not return them, that is itself the result of the test.
  • 4. Score three outcomes separately — right, wrong, and no answer. This is where most evaluations go wrong: collapsing "wrong" and "unknown" into one number hides the difference that matters. An unknown you can filter costs you a neutral greeting. A confident wrong answer costs you the customer.
  • 5. Now apply a threshold and score it again. Accept only the rows above your cut-off — say a probability of 0.9 with at least 50 samples behind it — and measure two things: what share of the file that covers, and the error rate within it. That pair is the real answer, because it tells you how much of the work disappears and how much risk comes with it.
  • 6. While you are in there, check the operational things. How long 300 rows take. What an error response looks like when you send rubbish. Whether the identical request twice returns the identical answer. Whether the trial needed a phone call.

Step five is why the response shape matters more than any headline figure: without a probability and a sample count there is no threshold to apply, so the honest answer to "how much of this file can I automate?" becomes "all of it or none of it". 100 free lookups a month covers a 300-row test three times over.

If you are also weighing up a language model

Increasingly the shortlist is not two APIs but an API and a prompt. That comparison turns on different things — cost per million rows, whether the same input gives the same output twice, and what you can show an auditor — so it has its own page.

What integrating actually looks like

Worth a look during evaluation rather than after it, because this is where an afternoon becomes a fortnight. One credit per lookup, 100 names per batch request, and no ceiling on how many requests you make.

  • Official clients for PHP, Python, Node, Java, Go, Ruby, Rust, Perl and .NET.
  • Full v2 reference with an OpenAPI description, plus RFC 7807 problem responses so errors are machine-readable.
  • A native Excel add-in and a Shopify app, ready-made integrations for Google Sheets, HubSpot, Salesforce, Zapier and more, and a hosted MCP server for AI tooling — the full list.
  • Bulk gender lookup for a one-off file, and nationality from a name if origin is part of what you need.

Pricing starts at €0.35 per 1,000 lookups and is on the pricing page in full — no quote needed to see it.

Frequently asked questions

What should I compare gender APIs on?

On what a single result hands you, not on a headline accuracy figure. A provider that returns a bare label forces you to trust every row or review every row; one that returns a probability and the number of records behind it lets you automate the confident rows and look at only the rest. After that: whether the same name can resolve differently by country, whether the answer is stable over time, what happens when the name is unknown, and where the data is processed.

How do I test one properly?

On your own list, not on a sample of famous names. Take a few hundred rows you already know the answer for — including the awkward ones — and run them through. 100 free lookups a month is enough for that, and any provider worth evaluating will let you do the same without a call.

Is a bigger name database better?

Not on its own. A name counts towards a total whether the database holds three records for it or fifty thousand, so the total cannot tell you which of your rows you can act on. The per-result evidence can. Size matters for how often you get any answer at all; the sample count matters for whether you should use it.

What accuracy do you publish?

No measured figure, deliberately. We have not run a published benchmark on a labelled name set, so quoting a percentage would be marketing rather than evidence. What we give you instead is the evidence per answer — a probability and a sample count on every result — so you can measure accuracy on your own data, which is the only number that describes your case. If a provider quotes you a single accuracy figure, ask which names it was measured on.

Does it matter where the data is processed?

It does if your list is customer names, because that is personal data. We are a German company, all servers are located in Germany, processing happens inside the EU, and a data-processing agreement can be requested in your account. Whoever you evaluate, ask for that in writing before the pilot rather than after it.

What are the questions people forget to ask?

Three. What happens to a name the database does not know — an empty field you can filter is worth more than a plausible guess you cannot spot. Whether the same request returns the same answer next year, which matters the moment an audit asks how a record was classified. And whether unused capacity expires.

RUN THE TEST ON YOUR OWN LIST

100 free lookups a month, no credit card and no call. Bring the rows you already know the answer to — including the awkward ones — and see what comes back beside each one.

Chat