A retro computer with a glowing Philippine sun on its screen.

A KOOYA RESEARCH PROJECT

A future
in our
own words.

Philippine languages. Everyday needs.
AI with a place for every Filipino.

Meet PinoSI

01 / THE WAY WE SPEAK

Kumusta,
future.

Tagalog. English. Ang Taglish sa pagitan.
AI na mas malapit sa buhay natin.

A greeting from us. Not a model demo.

02 / THE WORK STARTS HERE

Local words.
National ambition.

Kooya is adapting an open AI model for the way Filipinos speak, learn, and work.

A SMALL BEGINNING. A FILIPINO AMBITION.

03 / A FUTURE WE HELP BUILD

Para sa atin.
Para sa bukas.

For us. For the future.
The next chapter is ours to build.

Behind the ambition
RESEARCH IN PROGRESS
SCROLL TO EXPLORE

THE IDEA IS SIMPLE.

The future of AIshould have roomfor our voices.

Useful AI starts with understanding people. Their words. Their questions. Their everyday lives.

We’ll build on open research, test with native speakers, and let real results guide the next step.

WHAT WE’RE BUILDING

A model that
gets us.

PinoSI is Kooya’s research project to adapt an open AI model for the Philippines.

We start with an existing model and train it further on Philippine-language material. The goal: useful answers in Tagalog, English, Taglish, and more of our languages over time.

RESEARCH & DESIGN IN PROGRESS
  1. Start with a strong foundation.

    We’re building on an open model, rather than training from scratch. A small first round will help us find which training approach works best.

  2. Bring our languages into the training.

    The planned mix combines Philippine-language text and code. We check licences and sources, and use native-speaker-written Taglish questions to test the result.

  3. Make it useful. Check the work.

    Coding comes first because we can run the answers. We’ll measure whether the code works, whether Taglish improves, and whether English stays strong.

  4. Grow with evidence and partners.

    Future work could help people understand farm bulletins, public services, and learning materials. Each field needs suitable data, permissions, and its own tests before we expand.

THE FIRST EXPERIMENT / PLANNED

Does local language
training make a difference?

We’ll compare two versions of the same model on the same coding problems. Only one gets the extra Philippine-language training.

A

Open model Coding training

B

Open model Philippine-language training Coding training

Same questions. Same tests. Let the results decide.

Current facts. Sources you can check.

For future advisory and public-service tools, the plan is to read approved sources when answering, explain them clearly, and cite them. Changing facts should come from current documents, with human review where needed.

Explore the ten research areas

THE WORK BEHIND THE WORDS

Start small.
Prove it.

Ten research areas. A growing data foundation.
Now comes the work of training and testing.

10RESEARCH AREAS
3.16BLANGUAGE TOKENS · MEASURED
DATA INVENTORY · 11 OCT 2026
COLLECTED DATA · NOT MODEL RESULTS

DATA COLLECTED TO DATE 11 OCT 2026

01Philippine languages~110 GB
Language & Philippine knowledge · r3b base corpus

A foundation for language and Philippine knowledge. Tokens are the pieces of text an AI model learns from.

3.16B tokensMeasured · including 2.37B Tagalog tokens

IN THE INVENTORYfineweb2-fil, wiki-tl/regional, DepEd, seven web crawls, biblia, gutenberg-tl.

02Coding~254 GB
Coding · Athena

Our first pilot: test coding answers in English and Taglish by running the code.

~40–70B tokensEstimated · code

IN THE INVENTORYThe Stack v1, the-stack-dedup, OpenCodeInstruct, coder RawStore.

03Math & STEM~246 GB
Math · STEM

Math data for reasoning and problems with checkable answers. The inventory also includes evaluation sets.

~50B tokensApproximate

IN THE INVENTORYFineMath (149 GB), OpenWebMath (27 GB), NuminaMath CoT+1.5, gsm-symbolic; MATH, GSM8K and MMLU evaluations.

04Science~75 GB
Science · biology, chemistry & physics

Scientific questions and reasoning across biology, chemistry and physics. Synthetic means some material was generated by AI.

~8–12B tokensEstimated · much of the data is synthetic

IN THE INVENTORYNemotron SFT-Science-v2 (53 GB), RL-Science, OpenScienceReasoning-2, PubMedQA, ORD, physics-corpus, PaperSearchQA.

05Formal math~23 GB
Formal math · Lean

Mathematical proofs and tools for checking each logical step.

~3–5B tokensEstimated

IN THE INVENTORYNuminaMath-LEAN, Proof-Artifacts (15 GB), ntp-mathlib (7.8 GB), mathlib4, miniF2F.

06Agriculture~3.3 GB
Agriculture

Agricultural questions, plant images, and crop and price tables. The tables are structured data, rather than prose.

~0.3B tokensApproximate · plus structured PSA data

IN THE INVENTORYKisanVaani, CGIAR Q&A, PlantVillage images, PSA 23 crop + CPI tables.

07Supply chain~3.8 GB
Supply chain

Delivery, retail and demand datasets for practical logistics research.

Structured dataNo token total

IN THE INVENTORYLaDe delivery (3.6 GB), UCI Retail, M5, DataCo.

08Cybersecurity~30 GB
Cybersecurity · Argus

Security datasets for defensive research. Broader work remains subject to the project’s safety review.

~4–7B tokensEstimated · defensive scope

IN THE INVENTORYExisting RawStore (14 GB), PrimeVul (1.3 GB), CVEfixes (12.7 GB), MITRE, CVE, Sigma, MISP, YARA, OWASP and benchmarks.

09Biosafety & preparednessPending
Biosafety & preparedness

Public-health and preparedness sources. This part of the inventory is marked pending.

~0.1B tokensEstimated · less than 1 GB, pending

IN THE INVENTORYCDC NORS, WHO/OWID, BMBL and WHO manual, Nextstrain, epiparameter.

10Critical infrastructureTools & sims
Critical infrastructure

Power, water, transport and equipment simulation tools for infrastructure research.

Tools & simulatorsNot a text corpus

IN THE INVENTORYGrid2Op, pandapower, WNTR, SUMO simulators, OPSD, NASA PCoE.

Dataset inventory as of 11 October 2026. GB figures are approximate; token estimates are labelled. Collection does not mean every source is cleared for training or that a model has been trained. Licence, permission and evaluation checks still apply.

THE NEXT CHAPTER

$5MUSD · FUNDRAISING TARGET

A Philippine ambition.
A shared investment.
Explore a partnership ↗ (opens in a new tab)

US$50,000 invested by Kooya.
Our own funds. A start we believe in.

Research & designIN PROGRESS
Coding pilotFIRST PILOT PLANNED
EvaluationPLANNED

MADE HERE. FOR WHAT COMES NEXT.

Let’s build
what’s next.

Good questions. Local knowledge.
Let’s build this together.

Talk to Kooya (opens in a new tab)