Roblox announced on August 19, 2026 that it is contributing three updated open-source safety models to the Robust Open Online Safety Tools (ROOST) Model Community. The package includes version 2.0 of its PII Classifier, Roblox Sentinel version 2, and voice safety classifier version 3, along with a new evaluation dataset for PII detection in multiuser chat.

Versions of the models already run on Roblox to detect personal-information sharing, early child-endangerment signals, and voice-chat violations. Roblox joined ROOST as a founding member in early 2025 with partners including Google and OpenAI, and is sharing the tools so other platforms can train and tune their own moderation stacks.

What improved

PII Classifier 2.0 adds conversational context, expands language coverage from 17 to 189 languages, and raises the F1 score from 63.41 to 90.52 on Roblox’s reported benchmarks. The accompanying Roblox PII Classifier Benchmark uses synthetic multiuser chats that mimic phonetic bypasses, character substitution, and information split across turns.

Sentinel version 2 expands scoring combinations and evaluation options for early risk detection. Roblox said nearly 70% of cases it detected in the 12 months ended August 7, 2026, came from Sentinel’s early signals. Voice safety classifier version 3 covers 30 languages and eight violation categories, reaching 61% recall at a 1% false-positive rate across those languages after growing from about 95 million to 320 million parameters with distillation to keep latency usable.

Decoded Take

Open safety models only matter if competitors actually adopt and harden them. Roblox is buying industry goodwill while seeding evaluation data that matches the messy reality of gaming chat, not named-entity extraction demos. The metric to watch is whether ROOST peers publish comparable classifiers and shared red-team results, or whether this remains a one-way contribution that mainly burnishes Roblox’s trust narrative.