Every recording in our archives comes from a real, paid, consenting speaker. Here is exactly what that means in practice.
Recordings in our archives, including ALPIA, come from individuals who choose to take part as paid contributors. We do not scrape audio or text from the internet, social media or call recordings without explicit, individual consent.
Before recording anything, each contributor is told, in a language they understand:
Contributors are paid for their time and their recordings. Payment terms are agreed before recording begins and are not reduced or withheld based on how a recording is later used, within the terms the contributor originally agreed to.
We collect only the metadata that makes a recording scientifically or technically useful: language, variety or dialect, region, broad age band, and general recording conditions. We do not require full names, identification numbers or contact details to be attached to a public dataset entry.
We record how people actually speak, including regional and national varieties, code-switching and informal speech. This is treated as valuable data, not something to be corrected or standardised away.
Some datasets are collected specifically for a paying client and are not shared beyond that engagement unless the contributor agreed otherwise. Other datasets are collected for research or archive purposes and may be more widely available under the access tiers described on the ALPIA page. Where a dataset also supports our own internal research and development, this is disclosed to contributors as part of the consent process, not treated as a separate hidden use.
If you have contributed a recording and want to know how it is being used, request its removal, or ask a question about consent, contact us using the details below.
Email languages@grhoglobal.com with the subject line "Data and consent" and we will respond within a reasonable time.