A KOOYA RESEARCH PROJECT
A future
in our
own words.
Philippine languages. Everyday needs.
AI with a place for every Filipino.

A KOOYA RESEARCH PROJECT
Philippine languages. Everyday needs.
AI with a place for every Filipino.
01 / THE WAY WE SPEAK
Tagalog. English. Ang Taglish sa pagitan.
AI na mas malapit sa buhay natin.
02 / THE WORK STARTS HERE
Kooya is adapting an open AI model for the way Filipinos speak, learn, and work.
A SMALL BEGINNING. A FILIPINO AMBITION.03 / A FUTURE WE HELP BUILD
For us. For the future.
The next chapter is ours to build.
THE IDEA IS SIMPLE.
Useful AI starts with understanding people. Their words. Their questions. Their everyday lives.
We’ll build on open research, test with native speakers, and let real results guide the next step.
WHAT WE’RE BUILDING
PinoSI is Kooya’s research project to adapt an open AI model for the Philippines.
We start with an existing model and train it further on Philippine-language material. The goal: useful answers in Tagalog, English, Taglish, and more of our languages over time.
RESEARCH & DESIGN IN PROGRESSWe’re building on an open model, rather than training from scratch. A small first round will help us find which training approach works best.
The planned mix combines Philippine-language text and code. We check licences and sources, and use native-speaker-written Taglish questions to test the result.
Coding comes first because we can run the answers. We’ll measure whether the code works, whether Taglish improves, and whether English stays strong.
Future work could help people understand farm bulletins, public services, and learning materials. Each field needs suitable data, permissions, and its own tests before we expand.
THE FIRST EXPERIMENT / PLANNED
We’ll compare two versions of the same model on the same coding problems. Only one gets the extra Philippine-language training.
Open model Coding training
Open model Philippine-language training Coding training
Same questions. Same tests. Let the results decide.
For future advisory and public-service tools, the plan is to read approved sources when answering, explain them clearly, and cite them. Changing facts should come from current documents, with human review where needed.
THE WORK BEHIND THE WORDS
Ten research areas. A growing data foundation.
Now comes the work of training and testing.
DATA COLLECTED TO DATE 11 OCT 2026
A foundation for language and Philippine knowledge. Tokens are the pieces of text an AI model learns from.
IN THE INVENTORYfineweb2-fil, wiki-tl/regional, DepEd, seven web crawls, biblia, gutenberg-tl.
Our first pilot: test coding answers in English and Taglish by running the code.
IN THE INVENTORYThe Stack v1, the-stack-dedup, OpenCodeInstruct, coder RawStore.
Math data for reasoning and problems with checkable answers. The inventory also includes evaluation sets.
IN THE INVENTORYFineMath (149 GB), OpenWebMath (27 GB), NuminaMath CoT+1.5, gsm-symbolic; MATH, GSM8K and MMLU evaluations.
Scientific questions and reasoning across biology, chemistry and physics. Synthetic means some material was generated by AI.
IN THE INVENTORYNemotron SFT-Science-v2 (53 GB), RL-Science, OpenScienceReasoning-2, PubMedQA, ORD, physics-corpus, PaperSearchQA.
Mathematical proofs and tools for checking each logical step.
IN THE INVENTORYNuminaMath-LEAN, Proof-Artifacts (15 GB), ntp-mathlib (7.8 GB), mathlib4, miniF2F.
Agricultural questions, plant images, and crop and price tables. The tables are structured data, rather than prose.
IN THE INVENTORYKisanVaani, CGIAR Q&A, PlantVillage images, PSA 23 crop + CPI tables.
Delivery, retail and demand datasets for practical logistics research.
IN THE INVENTORYLaDe delivery (3.6 GB), UCI Retail, M5, DataCo.
Security datasets for defensive research. Broader work remains subject to the project’s safety review.
IN THE INVENTORYExisting RawStore (14 GB), PrimeVul (1.3 GB), CVEfixes (12.7 GB), MITRE, CVE, Sigma, MISP, YARA, OWASP and benchmarks.
Public-health and preparedness sources. This part of the inventory is marked pending.
IN THE INVENTORYCDC NORS, WHO/OWID, BMBL and WHO manual, Nextstrain, epiparameter.
Power, water, transport and equipment simulation tools for infrastructure research.
IN THE INVENTORYGrid2Op, pandapower, WNTR, SUMO simulators, OPSD, NASA PCoE.
Dataset inventory as of 11 October 2026. GB figures are approximate; token estimates are labelled. Collection does not mean every source is cleared for training or that a model has been trained. Licence, permission and evaluation checks still apply.
THE NEXT CHAPTER
A Philippine ambition.
A shared investment.
Explore a partnership ↗ (opens in a new tab)
US$50,000 invested by Kooya.
Our own funds. A start we believe in.
MADE HERE. FOR WHAT COMES NEXT.
Good questions. Local knowledge.
Let’s build this together.