What does source mean in V2?
The source field identifies the path that produced the final result: a stored dataset lookup, AI inference, or no selected source. It is not a citation to the original publisher of a dataset, a licensing certificate, or a guarantee that an inference is correct.
This guide describes the current V2 response contract and a first-party database profile. The database observations below are measured inventory and consistency checks; a reviewed upstream source catalog and an independent accuracy benchmark remain unpublished.
| source | What it tells you | What it does not establish |
|---|---|---|
dataset | A stored record was selected; an exact count tie can still return an unknown gender | The original publisher, collection method, licence, record age or independence of the observations |
ai | The final response came from the configured AI inference path; the gender can still be null | A counted sample, a dataset citation, or calibrated predictive accuracy |
none | No stored record was selected for the final dataset-only outcome | That the input has no possible name association in every population or language |
Database profile: 26 September 2026
Collection metadata was read at 2026-09-26T15:28:18.227Z; random samples were read at 2026-09-26T15:31:54.045Z. Counts come from MongoDB collStats metadata, not a separate full count. The general collection contains stored name records; the country collection contains stored name-country records. Their counts must not be added into a unique-name or unique-person total.
For each collection, MongoDB randomly sampled 5,000 documents and calculated the following aggregate checks on the database server. No raw names or records are included in the public report. These are first-party observations of a live database, not an independent audit or a frozen evaluation dataset.
| Observation | General records | Country-specific records |
|---|---|---|
| Collection | general | country-specific |
| Metadata document count | 2,483,220 | 4,080,692 |
| Stored unique key | name | name + country |
| Randomly sampled documents | 5,000 | 5,000 |
| Numeric rule failures in sample | 0 | 0 |
| male + female differs from total in sample | 0 | 0 |
| Equal male and female counts in sample | 124 | 94 |
| updated_at field present in sample | 667 / 5,000 | 741 / 5,000 |
What the database checks establish
The numeric check requires a positive safe-integer total and finite male/female counts between zero and total. A separate check compares male + female with total. All sampled counters passed those checks; this is not a claim that every record in the collections is consistent or that any prediction is correct.
Equal male/female counts occurred in 124 general and 94 country-specific sampled rows. These are ties in stored counts, not independently verified ambiguous identities or an API unknown rate. Actual API outcomes depend on inputs, Redis availability, candidate selection and AI settings.
A separate general-collection country-field inventory at 2026-09-26T15:33:13.811Z found 244 distinct stored values: 239 matched uppercase two-letter formatting and 5 did not. ISO membership, country-specific record distribution and supported-country coverage were not measured. These figures are not a count of supported countries.
No source or licence fields appeared in the two samples. This does not establish that provenance is absent from every record or from external systems. updated_at appeared as a numeric field in some records, but its values and meaning were not inspected; it does not establish freshness or an update schedule.
V2 serves dataset lookups from Redis. GenderAPI confirms that MongoDB and Redis contain the same dataset. The inventory above describes stored records; API answer coverage also depends on the submitted inputs, candidate matching and AI settings. The samples were unseeded, read without a frozen snapshot and had no independent reference labels; they cannot measure predictive accuracy, calibration or answer coverage.
Follow the evidence in the response
| Field | Meaning | Limit |
|---|---|---|
input | Echoes the input type, value and supplied country context | An input value is not independently verified identity information |
match.name / match.method | Shows the selected lookup candidate and whether it was normalized, tokenized, extracted as a substring, or inferred by a model | A token match does not guarantee that a true given name was isolated from a full name or handle |
match.scope / match.country | Shows whether the chosen stored row used country-specific or global lookup | A supplied country does not guarantee a country-specific row was available |
sample_count | The selected row's stored total; null for AI inference | Not necessarily unique people, independently sampled observations, or the size of an accuracy evaluation |
confidence / confidence_kind | Distinguishes a ratio from stored counts from a model-reported score | Neither value is a verified probability of an individual's gender identity |
country / country_source | The returned record's country metadata or an AI name association, when available | Not nationality, residence, verified origin, or proof that the input country was matched |
How input matching affects a result
Dataset lookup normalizes text and derives candidate tokens. For email inputs it derives candidates from the local part; for usernames it removes an optional leading @. Ordinary email and username lookup can also try bounded substring candidates. If multiple candidate records are available, the record with the largest stored total is selected. This can select a token other than the intended given name.
When country context is supplied, the lookup tries the country row for each candidate and uses its global row when the country row is absent. The final result reports the selected match. A country-specific row that is ambiguous is not automatically replaced by a global row for that candidate.
The normal AI prompt asks for a recognizable given name. forceToGenderize also allows personal nickname or alias associations, so a non-null gender with name: null is valid in that mode. These are different inference tasks and need separate evaluation.
Dataset lookup and AI processing
When AI is invoked, V2 sends the submitted type, value and country context to the configured AI inference service. For email inputs, the prompt tells the model to ignore the domain and plus-addressing suffix; this is not a promise that those characters are removed from the value before AI processing. Use ai_mode: off without forceToGenderize when your workflow requires dataset-only prediction.
AI response validation checks JSON fields, types and cross-field consistency. It does not prove that a returned name, gender association or country association is factually correct. The public result does not include the model identifier, model version or a citation for its inference.
| Option | Inference path | Evidence returned |
|---|---|---|
| ai_mode: off | Uses dataset lookup only | Selected dataset metadata, or an unknown result with no selected record |
| ai_mode: fallback | Tries the dataset first; invokes AI if no gender is returned | The final path's evidence, not a combined dataset-plus-AI sample |
| ai_mode: always | Uses the configured AI inference path directly | Model-reported score or an unknown result with null confidence; sample_count is always null |
| forceToGenderize: true | Tries the dataset, then nickname-capable AI when unresolved | Dataset evidence or model-reported alias association; name may be null |
Source documentation that remains unpublished
A reviewed source catalog is not published here. The response contract and implementation behavior alone cannot verify data origin, permission to use a source, representativeness, or the quality of its labels. Before making source-specific claims, the following evidence needs to be documented and reviewed.
This documentation review does not establish that upstream sources are licensed, unlicensed, complete or incomplete. It states the limits of the public evidence. Contact GenderAPI when your intended use requires a source-specific assessment.
| Evidence | Required record | Public status |
|---|---|---|
| Origin and permitted use | Named source or publisher, collection period, licence or other documented basis for use | No reviewed catalog published here |
| Processing lineage | Transformation, normalization, deduplication, aggregation and any synthetic or model-generated contribution | Not exposed by a source: dataset label |
| Population coverage | Country and script coverage, selection process, label definitions and documented bias | No independently evaluated coverage figures published here |
| Freshness | Snapshot identifiers, per-source update dates and a supported update schedule | No per-record freshness or snapshot identifier in the V2 response |
| Quality and correction | Validation results, overlap checks, correction/removal handling and review responsibility | No source-specific audit results published here |
Privacy, processing and responsible use
Source provenance, personal-data handling and legal terms are different questions. Consult the current published policies for processing roles, retention, providers and contractual terms. This technical guide does not add a new retention promise or certify legal compliance.
Send only the information needed for your workflow. Preserve confidence kinds and unknown outcomes, and do not describe an inference as a person's self-declared identity. Do not use inferred gender as the basis for consequential decisions about an individual.
Accuracy evidence and review history
A source label explains the inference path, not its measured quality. Use the accuracy guide to distinguish response scores, answer coverage, errors and the information needed for a reproducible evaluation.
| Date | Documentation revision | Change |
|---|---|---|
| 2026-09-26 | 1.0.1 | Published a dated MongoDB inventory and bounded sample consistency checks alongside V2 source and match fields, AI processing boundaries, and remaining source-catalog gaps. Recorded GenderAPI’s confirmation that MongoDB and Redis contain the same dataset. No accuracy percentage or source-licence claim is introduced. |