Field notes

6 min read

Millions of sentences we never wrote

Millions of sentences we never wrote

Why we put a rule engine next to our CMS, and when you might want one too. The product below is fictionalized. The problem and the architecture are real.

Why we put a rule engine next to our CMS, and when you might want one too. The product below is fictionalized. The problem and the architecture are real.

BY Joren Van Hemelrijk

Server-side Engineer

Joren is a serverside engineer at November Five. He focuses on building performant, scalable and reliable backend systems using cloud infrastructure. Products that have to deal with quick bursts of high amounts of traffic are a challenge he will never shy away from.

TL;DR

TL;DR

  • When to use a rule engine for your content: when the text you show describes the user's data (prices, limits, conditions) and the number of special cases keeps growing.

  • The problem: writing millions of these sentences by hand is impossible, and a copy in the CMS will eventually drift from the real data and make false promises to customers.

  • Why not just code? It works at first, but every new exception means a developer ticket and a deploy, while the people who actually know the rules can't touch them.

  • The fix: a rule engine turns that logic into an editable table, so non-developers can change rules themselves and every sentence is built fresh from live data, in every language.

  • When you don't need one: if your text is static, use a CMS. If your rules are few and stable, simple code is fine.

The problem

Picture the app of a mobile carrier. Before a trip, you open it, pick a country, and it tells you what your plan does there:

Data works like at home, up to 25 GB. Calls cost €0.49 per minute. No coverage included. Turn on a day pass first.

Simple screen. The catch is in the numbers: about 200 destinations, about 3,000 plan variants (every consumer, business, and legacy plan sold in the last decade), and every sentence in three languages. That multiplies out to millions of sentences.

And every one of them has to be right, because this is money. Tell someone data is included when their plan charges per megabyte, and they come home to a €900 bill and a screenshot of your app saying otherwise. You don’t want your call center to be overwhelmed by unhappy customers because of these kind of mistakes.

You can’t write them by hand

First idea: it’s text, so it goes in the CMS like all our other text.

Volume aside, there’s a worse problem. “Data works like at home, up to 25 GB” is not copy. It’s a claim about data. The real answer lives in the structured config per plan and destination:

{
   "destination": "JP",
   "dataIncluded": false,
   "ratePerMB": 0.12,
   "dayPass": 7.99
}
{
   "destination": "JP",
   "dataIncluded": false,
   "ratePerMB": 0.12,
   "dayPass": 7.99
}
{
   "destination": "JP",
   "dataIncluded": false,
   "ratePerMB": 0.12,
   "dayPass": 7.99
}

The moment you write that sentence down somewhere else, you have two sources of truth. They will drift. The pricing team changes a country’s rate, nobody updates the CMS record, and a user gets a promise their plan won’t keep.

The rule we landed on: if a sentence describes data, derive it from the data. Never author it.

Fine, derive it in code

Config in, sentence out:

if (d.dataIncluded)        return i18n("like_home", { cap: d.fairUseGB });
if (d.dayPass != null)     return i18n("day_pass", { price: d.dayPass });
return i18n("per_mb", { rate: d.ratePerMB });
if (d.dataIncluded)        return i18n("like_home", { cap: d.fairUseGB });
if (d.dayPass != null)     return i18n("day_pass", { price: d.dayPass });
return i18n("per_mb", { rate: d.ratePerMB });
if (d.dataIncluded)        return i18n("like_home", { cap: d.fairUseGB });
if (d.dayPass != null)     return i18n("day_pass", { price: d.dayPass });
return i18n("per_mb", { rate: d.ratePerMB });

This works. The strings live in the translation platform, the code picks the right one, nothing can drift. If your rules look like this and stay like this, stop here. You don’t need a rule engine.

Ours didn’t stay like this. The variations kept coming, and none of them were copy changes:

  • “After the fair-use cap we throttle to 512 kbps.”

  • “Business plans pool data across the whole team.”

  • “Cruise ships and planes are satellite rates even inside the EU.”

  • “Some legacy plans have a day pass and a per-MB rate.”

Each one is a new condition. A new if, a ticket, a review, a deploy. The pricing team, the people who actually know these rules, couldn’t touch them. They could only file tickets and wait for us to translate their knowledge into code. We were the bottleneck for logic we understood worst.

Remember from the intro: 200 destinations x 3,000+ plan variants x 3 languages ≈ 2 million variations

Move the ifs out of the code

So we did to the logic what everyone already does to the strings: moved it out of the codebase.

A rule engine lets you write that if-chain as a decision table:

dataIncluded

fairUseGB

dayPass

ratePerMB

→ sentence

true

null

*

*

“Data works
like at home.”

true

> 0

*

*

“Data works like
at home, up to
{fairUseGB} GB.”

false

*

> 0

*

“No data included.
A day pass costs
€{dayPass}.”

false

*

null

> 0

“Data costs
€{ratePerMB}
per MB.”

Read top to bottom, first match wins. Each row’s output holds the sentence in all three languages, so a translation can never end up attached to the wrong rule.

  1. The pricing team edits this themselves. A new condition is a new column, a new case is a new row. They test a rule against a real plan in the built-in simulator, then publish. The backend picks up the new rules within seconds. No ticket, no deploy.

  2. The table is the spec. When someone asks “what do we tell a business user on a cruise ship?”, nobody digs through code. You read the row.

  3. String interpolation can collapse thousands of different sentences using a handful of rules.

We used GoRules. It has a very user friendly Business Rules Management System (BRMS) where the rules live and get authored. This outputs optimized JSON Decision Model (JDM) files which are consumed by a fast stateless (and open source) engine to evaluate them. Other engines exist; the shape matters more than the product.

How it runs in production

  1. The backend fetches the user’s plan config for that destination from the billing system.

  2. It sends that config to the rule engine.

  3. The engine returns one sentence per service (data, calls, texts), in three languages.

  4. The backend merges that with the normal CMS content (country names, flags, help articles) into one response for the app.

// in
{
   "destination": "JP",
   "dataIncluded": false,
   "ratePerMB": null,
   "dayPass": 7.99
}

// out
{
  "destination": "JP",
  "data": {
    "en": "No data included. A day pass costs €7.99.",
    "fr": "Pas de données incluses. Un pass journalier coûte 7,99 €.",
    "de": "Keine Daten inklusive. Ein Tagespass kostet 7,99 €."
  }
}
// in
{
   "destination": "JP",
   "dataIncluded": false,
   "ratePerMB": null,
   "dayPass": 7.99
}

// out
{
  "destination": "JP",
  "data": {
    "en": "No data included. A day pass costs €7.99.",
    "fr": "Pas de données incluses. Un pass journalier coûte 7,99 €.",
    "de": "Keine Daten inklusive. Ein Tagespass kostet 7,99 €."
  }
}
// in
{
   "destination": "JP",
   "dataIncluded": false,
   "ratePerMB": null,
   "dayPass": 7.99
}

// out
{
  "destination": "JP",
  "data": {
    "en": "No data included. A day pass costs €7.99.",
    "fr": "Pas de données incluses. Un pass journalier coûte 7,99 €.",
    "de": "Keine Daten inklusive. Ein Tagespass kostet 7,99 €."
  }
}

Nothing is stored or pre-computed. Every sentence is built from the live plan config on every request, so there is no second version of the truth to drift.

When you’d want this

You don’t need a rule engine for all situations. You might if all of these sound familiar:

  • The text you show is a description of the user’s data, not static copy.

  • The distinctions keep growing. Every month someone needs the text to say one more thing in one more situation.

  • The people who know the rules don’t deploy code.

  • You’re multilingual, and translations have to stay glued to the rule they describe.

If the text is static, use a CMS. If the rules are few and stable, an if-chain with translation keys is fine. The rule engine earns its place when the logic changes as often as the copy, and the people changing it aren’t developers.

We shipped millions of sentences. We wrote about 100 rules.

NOVEMBER FIVE

If the product is core to your business, you can't afford to get the team wrong.

NOVEMBER FIVE

If the product is core to your business, you can't afford to get the team wrong.

NOVEMBER FIVE

If the product is core to your business, you can't afford to get the team wrong.