In this post
The industry default for moderating user content is reactive: publish first, report later, remove if someone complains.
It works at scale, and it fails exactly in the first few hours — which are the hours when content circulates. By the time removal happens, the damage in a serious case is done, and the victim did the work of reporting.
In EuRi, a meme social network, we inverted that. This post is about what the inversion cost.
How it works
user creates a post
-> Cloud Function
-> Gemini analyses
-> passed? yes -> goes live
no -> does not
There is no window between publishing and moderating, because there is no publishing before moderation.
Choosing a Cloud Function rather than doing this in the app is the point that matters most and gets discussed least: client-side moderation is optional moderation. A modified app skips the check and writes straight to the database.
And here it pays to be precise, because this is where many people believe they have solved it and have not: the Cloud Function is not what guarantees anything. The Function is just a path; if the database accepts writes directly from the app, the path is bypassable. The guarantee lives in the database security rules, which must deny client writes to the feed collection and allow them only from the backend.
| Where the guarantee lives | What a modified app does |
|---|---|
| Only in the Function | writes straight to the database, skipping moderation |
| In the database security rules | the write is refused, and there is nothing to bypass |
Writing the Function is the visible work. Closing the side door is the work that makes the Function worth anything.
The detail that changes everything: the text is inside the image
A meme is not text with an image attached. Most of the time it is an image with the text drawn into it — and the text is where nearly all the problematic content lives.
That invalidates the intuitive solution. A classifier reading the post caption analyses precisely the part that does not matter: the author writes "lol look at this" in the caption while the slur is rendered in white across the top of the image, where no text filter can see it.
The analysis has to cover the whole image, with the model reading what is written on it and judging text and picture together — because separately they are usually harmless, and the meaning comes from the combination. That detail is what makes this problem more expensive than "call a moderation API".
What cost more than expected
Latency on the critical path
Calling a model before publishing adds waiting, and the wait is variable.
The real problem is not the time: it is perception. A button that goes quiet for two seconds reads as a freeze. And the obvious answer — a loading spinner — solves it badly, because it explains nothing.
| Approach | Perception |
|---|---|
| Frozen button, no feedback | broken app |
| Generic spinner | slowness |
| "Checking your post", explicit state | a process, with a visible end |
The third also creates the opportunity to explain a rejection when it happens — the difference between a user who understands the rule and one who thinks the app is arbitrary.
False positives
This is where the most time went, and the reason is intrinsic to the content: memes are ambiguous by construction.
Irony, sarcasm, self-deprecation and in-group humour look like aggression to a naive classifier. A keyword filter blocks half of the legitimate material, and a humour app with a filter like that is not a safe humour app — it is a broken one.
Early prompt versions erred on the restrictive side unacceptably often. What fixed it was not tuning severity: it was replacing description with examples.
A prompt that describes offence in the abstract produces an anxious classifier. A prompt with concrete cases — including things that should pass — behaves far better. The list of passing cases matters as much as the blocking one.
AI moderation is not a switch. It is a parameter you calibrate against your own content — and erring on the restrictive side costs you the product.
There is a second-order effect worth planning for: a rejected post is a moment of friction with someone who did nothing wrong. Handled badly — a generic "this violates our guidelines" — it reads as an accusation, and the user either leaves or reposts the same thing three more times. Handled well, with a specific reason and a way to appeal, it becomes the clearest signal the product ever sends about what kind of place it intends to be.
Which means the rejection copy is not a detail of the moderation system. On a humour app, it is the moderation system's user interface, and it will be read far more often than any policy page.
The failure decision
Any synchronous moderation architecture must answer one question before shipping:
If the analysis service is unavailable, does the post publish or does it hold?
Failing open keeps the app working and opens an unfiltered window precisely during an outage — when nobody is watching.
Failing closed holds content and frustrates legitimate users.
We chose to fail closed, with a visible "under review" state. The reason is specific to this product: if the differentiator is moderation, a period without moderation is not degradation — it is the absence of the thing we promised.
In a product where moderation is incidental, the opposite choice would be reasonable. What matters is that the decision is made deliberately, not discovered during the first incident.
What the inversion actually solves
It removes the circulation window. Serious content does not sit live for twenty minutes.
It takes the burden off the victim. In the reactive model, the system depends on someone being harmed and having the energy to report it. That design outsources moderation work to exactly the wrong person.
It reduces human volume. What reaches manual review is the edge case, not the obvious one.
What it does not solve
Reports still matter. Borderline cases, context and disputes still need people.
Context-only harm. A harmless image aimed at one specific person can be harassment, and no classifier looking at the post in isolation sees that. This is the hardest gap and it is structural: the unit of analysis is a post, and the unit of harm is a relationship between two accounts over time. Closing it needs signals the moderation call never receives.
Behaviour change. Users learn what passes. A filter that is not revisited becomes a filter that is routed around. The adversary here is not a hacker; it is an ordinary user with time and a sense of humour, and they iterate faster than any release cycle.
It is not free. Every post costs a model call, including the rejected ones. That inverts a comfortable intuition: in an ordinary social app a heavy poster is a valuable user; here they are also a line on the invoice. And anyone trying to flood the app with prohibited content spends your money while being blocked, because the rejection only exists after the analysis.
The practical consequence is that a per-account posting limit stops being hygiene and becomes part of the cost architecture. It is not protection against spam in the feed — the filter already handles that. It is protection against spam on the service bill.
Frequently asked questions
How much latency does this add?
Enough for users to notice. Calling a model on the critical path of publishing is not free, and the response time is not constant. The work was not to eliminate the wait — it was to make it legible in the UI so it does not read as a freeze.
What happens if the model is unavailable?
That is the most important design decision in the system: fail open (publish) or fail closed (hold). We hold and show an 'under review' state, because in an app whose differentiator is moderation, publishing unfiltered during an outage destroys the promise itself.
How do you tune the filter without killing the humour?
With examples, not adjectives. A prompt describing offence in the abstract produces an anxious classifier. A prompt with concrete cases of what passes and what does not — including irony and self-deprecation — behaves far better.
Does this replace human moderation?
No. It reduces the volume reaching a human and removes the window in which the worst content circulates. Reports, review and edge-case judgement still need people.