Federal grant · project grant (b)
Collaborative Research: Redddot Phase 2: Inclusive American Language Technologies -more Than 350 Languages and Many Additional Variants and Dialects Are Spoken in the United States and Yet, Voice Technology Recognizes Only a Handful. This Research Will Create Crucial Training Datasets, Predominantly Optimized for Speech Recognition (speech-to-text), for Three Underrepresented, American Sociolinguistic Contexts ? a Sociolect, a Code-switching Language Context, and an Indigenous Language. the Methodology for Co-creating These Datasets With Communities Prioritizes Building the Agency, Skills, and Knowledge Required for People to Use and Apply Their Dataset to Serve Their Own Social and Economic Context. Inclusive Speech-to-text Technology That Recognizes More American Language Dialects Means That More Americans Can Access Critical Information Across Citizen Services, Finance, Education, Health, and Justice. the Project Iterates a Community-mobilizing, Inherently Capacity-building, Applied Methodology for Creating Crucial Machine-learning Datasets, Predominantly Optimized for Speech Recognition (speech-to-text). the Data Creation Process (text and Audio) for These Datasets Will Be Run, Hosted, and Released Through an Open-source Platform and Infrastructure to Ensure Public Accessibility. Communities Will Co-create the Datasets From Design Phase to Quality Assurance, With Space to Shape the Governance Framework, Diversity Criteria, and Domain Representation. This Program Will: (1) Bridge Critical Gaps for Innovative Technological Research on Under-represented Languages and Variants; (2) Evolve Understanding of Culturally-conscious, Consent-centric Modes of Community Participation in the Building of Artificial Intelligence (ai); and (3) Accelerate First-language Language Technology Tooling in Key Economic Domains Such as Health, Education, Justice, and Agriculture, Thereby Accelerating Pathways to Societal and Economic Benefits. the Project Will Also Advance Skills Development in Machine Learning by Actively Involving Individuals Who Speak These Underrepresented Language Variants in the Data Collection Process. the Project Methodology Is Applied Pedagogy, Through Teaching Communities About Ai Training Datasets by Involving Them in Their Design and Build. This Skill-building Approach Can Lead to Improved Community Representation Within Stem Professions, as Well as Immediately Mitigating Dataset Biases and Potential Harms. This Award Reflects NSF'S Statutory Mission and Has Been Deemed Worthy of Support Through Evaluation Using the Foundation's Intellectual Merit and Broader Impacts Review Criteria.- Subawards Are Not Planned for This Award.
Committed
$1.1 Million
Paid out
$9.3K
<1%
Committed, not yet paid
$1.0M
99%
Loading…
Everything here is this single award's whole record — signed, amended, paid — not a fiscal-year slice. The by-year charts elsewhere split an award across the years it was committed; this page keeps it whole.
Committed is what the government has legally promised on this award so far. Contracts can also carry a ceiling — the maximum if every option is exercised. Unspent ceiling is headroom, not money owed.
The cash actually disbursed against this award. The gap from committed is the disbursement pipeline: promised, not yet cashed.
Each transaction is a signing event — an action that created or changed the award, dated the day it was signed — not a payment. Negative amounts are real: money de-committed at closeout or renegotiation.
One bar, the award’s whole arithmetic: paid out, then committed, not yet paid, then unspent ceiling.