EverestQ

Multilingual AI infrastructure focused on low-resource South Asian languages including Nepali, Maithili, Bhojpuri and related languages.

Overview

EverestQ is a multilingual large language model focused on low-resource South Asian languages. Built under Optumina Research Labs, the project aims to create high-quality language models for languages that are underrepresented in current AI systems.

Why EverestQ exists

Most large language models are trained on predominantly English and high-resource language data. Languages like Nepali, Maithili, and Bhojpuri — spoken by millions — have very limited representation in current AI systems. EverestQ exists to address this gap.

The language problem

Low-resource languages face unique challenges: limited training data, lack of standardized benchmarks, and minimal research attention. Building models for these languages requires different approaches to data collection, training, and evaluation.

Model architecture

everestq ~ $ architecture
Transformer-based
Multilingual pretraining
Low-resource optimized

Training & Data

Training involves curating datasets from diverse sources for South Asian languages, with careful attention to quality, diversity, and cultural context.

Infrastructure

Built on modern ML infrastructure with a focus on efficient training and inference for multilingual models.

Current status

status: Building