Back to the blog
AI & Moderation

Moderation Just Got a Decision Model That Reads Your Policy in 35 Milliseconds. The Threshold Is Now the Policy.

October 7, 2026 6 min read

On Tuesday, October 6, a company called Musubi released PolicyLM-1.7B, an open-weights model under the Apache 2.0 license, built for one job: content moderation. You give it a message and your own content policy, written as short rules in plain language, and it returns a score between 0 and 1 for each category of that policy. It generates no text. On a single 24 GB NVIDIA L4, Musubi measures a median of 35 milliseconds per short chat message with up to six categories.

TechCrunch framed the launch as decision models reaching moderation. The category is three weeks old. TypeSafe launched Jev on September 15. OpenAI announced a Decisions API at DevDay on September 29, returning answers developers can use to "classify content, route requests, or choose an agent's next action." Amazon's Strands Labs released an open-source Strands Decider 2B on October 1. Musubi does not hide the lineage: "If Jev caught your eye, PolicyLM-1.7B is the same kind of model, trained specifically for content moderation, that you can run yourself."

I use Jev daily for small yes-or-no calls, like checking the claims in my own posts. I do not moderate a platform. But moderation is where answering with a number instead of a verdict matters most.

The biggest decision problem on the internet

On October 7, the EU's DSA transparency database, where online platforms file the "statement of reasons" they owe a user each time they remove or restrict that user's content, showed 4,301,448,066 statements submitted over the last 180 days by 374 platforms. Forty percent were fully automated decisions.

Each statement carries a required field saying whether the decision was fully automated, partially automated or not automated. None of the documented fields records how sure the machine was. The European Commission describes the statement as a tool to help users "understand and potentially challenge content moderation decisions."

Until now, the automated part has mostly run on two tools. Fixed classifiers are fast and cheap, but they score their own built-in categories, so changing a rule means relabeling data and retraining. Large language models read your actual policy, but in Musubi's comparison they usually take hundreds of milliseconds or more per check, too slow and too expensive for live chat.

A decision model sits between the two. It reads the policy like an LLM and answers at classifier speed. The interesting part is not the speed. It is that the answer is a number.

What changes when the decider is a number

Thresholds become per context. PolicyLM ships two presets: "precision", the default, at 0.335, for surfaces where violations are rare such as live chat, and "balanced" at 0.275, for queues where a miss costs more than a false flag. Each category can get its own cutoff. TypeSafe's docs agree: "A confidence threshold is not one number." A username, a DM and a marketplace listing do not deserve the same cutoff, and a policy team can now move it without a retraining cycle. The threshold is the policy, written in the only language the model obeys.

Under the default preset, a score of 0.34 flags a message. That is a score, not a promise that a third of such messages break the rules. TypeSafe defines calibration as outcomes given 0.8 happening about 80 percent of the time "across many predictions." Whether a model's numbers mean that on your traffic is for you to measure; Musubi itself says to calibrate on your own content before going live.

Appeals can replay a decision, if you kept it. An appeal asks one question: was this removal right under the rules in force that day? If you stored the score, the category, the policy version, the threshold and the model revision, you can answer exactly, and say whether it would pass today. I argued for that pattern in Store the Probability, Never the Verdict. Moderation is where it pays the most.

Do not replay by rerunning the model. PolicyLM's model card notes that bfloat16 reads on CUDA and on Apple silicon differed by up to 0.083, and that CPU float32 changed a few decisions; for exact, repeatable scores it recommends float32. TypeSafe ran a moderation rubric over one borderline post 15 times: Jev kept its top label 90.8 percent of the time and flipped on 2 of 8 questions. Its docs add that a threshold "does not make the model deterministic." The log is the record. The model is not.

Auditability gets more honest. A decision model writes no reason, and Musubi lists "No reasons provided" among its limitations. For statements of reasons, I think that is closer to a feature. A record such as "off-platform trading, policy version 14, score 0.71, threshold 0.335" is a reason anyone can check. A rationale generated after the fact may not describe how the verdict was reached.

The cost per decision collapses. Jev costs $42 per billion input tokens, and outputs are free. A 2,000-token call, policy plus message, costs $0.000084, so a million decisions cost $84. PolicyLM's card says one L4 sustained 34.4 messages per second at a p95 of 150 milliseconds or less; that is almost three million messages a day on one card. Musubi's pitch is "score all of your traffic instead of sampling it." Moderation moves from sampling to census.

Where it still breaks

Adversarial users. The model card lists "a security boundary against adversarial users" as out of scope. The helper's text cleaner undoes look-alike characters, spaced-out letters and base64, and caught 10 to 12 points more disguised violations on Musubi's test set, but "leetspeak mostly gets through." People trying to get past a filter iterate faster than any policy team.

Drift. The card warns that scores shift "with policy wording, language, device and numeric format." Policy edits are not free either: on Musubi's benchmark, about 1 in 2 single-clause edits actually changed the decision. Hosted models move too: Jev's jev-latest alias currently points to jev-1.13.0, and each response names the version that answered. And read the card next to the blog post. The post says a custom fine-tuned PolicyLM runs on a platform handling more than a million messages a day. The card for the released weights says "Not yet tested on live traffic."

Accuracy is still bought with time. On Musubi's custom-policy benchmark, OpenAI's open-weight gpt-oss-safeguard-20B scores 0.909 accuracy against PolicyLM's 0.842 and follows 0.716 of policy edits against 0.528, at a median of 349 milliseconds per message against 22, both on one H100. Musubi's own comparison sends appeals, bans and takedowns to larger models.

The human queue. In TypeSafe's borderline test, a 0.60 floor raised agreement to 99.2 percent, but only 74.2 percent of answers were automatic; the rest went to human review. That is one hard post, not a rate. The lesson holds anyway: a cheap decider does not remove the queue. Your thresholds decide how big it is, so set them with your reviewer headcount on the table.

What I would do on Monday

Log every automated decision as a row: score, category, policy version, threshold in force, model revision, numeric format. Set thresholds per surface, not per platform. Define an explicit uncertain band and track its daily volume against what your reviewers can clear. Build a labeled sample from your own traffic and check monthly that messages scored around 0.7 are violations about as often as you assumed. Pin model versions. Answer appeals from the log.

Then ask who runs the model. In the LegalTech I build, client documents are sensitive, so where a model runs is a question I ask before accuracy. For moderation of private messages, open weights on your own infrastructure are not a detail.

The model gives you a number. What the number means is still a policy decision, and it now deserves the same review as the policy text.

Sources

A project like this one?

I design and deploy products like this. Let's talk.

Let's talk