The top-up nobody knows went through
A customer asks for a prepaid airtime top-up at a corner store. The system debits the store’s balance and sends the request to the carrier. The carrier does not answer.
Did the top-up reach the phone? At that moment nobody knows. And with the carrier, nobody can ask: the answer arrives tomorrow, in the reconciliation.
For years I worked on an electronic top-up platform. Requests came in from terminals, apps and points of sale, and reached the carrier after two, three and up to four hops, each with its own protocol and its own timeout. Everything happened: double top-ups, charges with no top-up. This note is about why, and about the state almost no system designs in from the start.
The figures in the examples are illustrative, calculated to show the mechanism. They are not that platform’s numbers.
There are three answers, not two
A system that moves money is usually thought of with two outcomes: the operation went through, or it did not. On a network there is a third: unknown. The request went out and the answer never came back. The request may have been lost, and the top-up does not exist. Or the answer may have been lost, and the top-up is already on the phone.
From this side of the wire, the two cases look identical.
Every hop is one more place where an answer can get lost. If each one loses one in two thousand, with one hop that is 50 top-ups in doubt out of every hundred thousand; with four hops it is 200. Every day. It is not a bug you fix once: it is a probability you manage forever.
And the hardest case is at the end of the chain: the carrier does not answer, or its answer gets lost, but the internal system has already debited.
Everyone decides what to do with “unknown”
When the screen hangs, somebody decides, and every decision is a bet with someone else’s money:
- The cashier presses the button again. If the first one did go through, the customer gets two top-ups and the store pays for two.
- The terminal retries on its own. Same thing, without anyone noticing.
- The platform retries toward the carrier. Same thing, one hop deeper.
- The platform marks the operation as failed and refunds the store. If the carrier did apply it, the platform gave away a top-up.
- The platform marks it as good. If the carrier did not apply it, the customer paid and got nothing, and is going to complain.
None of them is right. They are all bets as long as nobody knows what happened.
An ID that travels end to end
The first defense is for every top-up to have a unique identifier from the moment it is born, in the terminal or the app, and for that same identifier to travel through every hop. We called it the externalID.
With it, a retry stops being a new top-up. If a request arrives with an externalID that was already processed, the system does not run it again: it answers with the result of the first one. The cashier can press ten times and the store pays once.
And it makes it possible to ask. Backwards, from the platform toward the points of sale, you could: the terminal could ask what happened to its externalID instead of sending another top-up.
What you cannot ask
With the carriers, you could not. There was no status query. The only thing was batch reconciliation: a file every 24 hours with what the carrier says it applied, against what the platform says it sent.
That means a top-up in doubt stays in doubt until the end of the day, or until the next day. Meanwhile, the store’s balance has already been debited, the customer may be complaining, and the platform has nothing to tell them.
And reconciliation, when it comes, is torture: cross-matching files, finding the differences, deciding each one, and sometimes resolving them by hand.
Why not reverse
The way out that looks obvious is the reversal: when in doubt, cancel the operation with the carrier and return the money. In this business reversals are not desirable: they kill the business.
The margin on a top-up is very small. Every reversal costs work, time and arguments with the carrier, and eats the profit of many sales. And if you reverse something that did go through, the loss is total.
What I would do from the start
Design “unknown” as a real state. In the database, on the cashier’s screen and in the reports: in progress, not failed or successful. A screen that says “in progress, do not repeat it” is worth more than any later validation.
The externalID from the origin, mandatory, and honored at every hop. A request with no identifier does not get in. A repeated identifier returns the stored result and never runs again.
No human retries blindly. The retry button checks the status first. If there is no answer, it shows that the operation is in progress; it does not send another.
Staggered timeouts. Each hop has to wait longer than the next one. If the terminal gives up before the platform does, the terminal retries while the first request is still alive inside.
Automatic reconciliation from day one. Not as a spreadsheet someone puts together once there are problems, but as part of the system: the carrier’s file comes in, gets matched on its own, and a person only gets the differences.
Measure the doubts. How many top-ups stay in progress, per carrier, per hour and per hop. When that number goes up, something broke, and it is better to know today than in tomorrow’s reconciliation.
None of these measures eliminates “unknown”. What they do is stop it from being a bet: the money sits in a state everyone can see, nobody duplicates it out of desperation, and reconciliation only has to resolve what truly could not be known. Mitigating it was a lot of programming and infrastructure work. There was no shortcut.
Israel Negrete Lepe is an electronics engineer. Twenty-five years building systems that make it to production: telemetry, vehicle tracking, municipal video surveillance, RFID asset control and transactional platforms.