Talk
Building an emergency software
A keynote about four Elixir systems built out of Kisumu and about the friend whose death started the first one.
- Event
- ElixirConf EU Virtual 2020
- Date
- 7 to 8 October 2020
- Location
- Online
- Length
- 49 minutes
- Keynote
- Elixir
- Phoenix Channels
- USSD
- Offline-first
- Non-profit
Takeaways
I gave this keynote with Linda Achieng, a software engineer at Syncro and a friend of about a decade. We set expectations early: no fancy technical terms, no framework tour. Just what we had built and why.
The brief came from a funeral, not a market study
Nailinda exists because a friend died in traffic on the way to the only hospital we knew in a city we did not know. His mother said at the eulogy that he would have lived if we had got there earlier. That sentence is the product requirement.
It runs on USSD because the people who need it do not have data
Anyone who can open Google Maps can already find a hospital. The users we built for dial a short code, have no smartphone and sometimes have no airtime left. That decision put a mobile app near the bottom of the list, not the top.
We picked Elixir for the connection, not the syntax
An emergency service cannot lose a socket and shrug. Reconnection had to be something we did not think about every day.
Free was a design decision, not a missing business model
If the hospital pays, the hospital gets reluctant. If the patient pays, it stops being an emergency and starts being a transaction. So we absorbed the cost, which meant the runtime had to be cheap to run at scale.
One rewrite took a monthly bill from 15,000 dollars to the price of servers
We were renting real-time infrastructure. Past 100 websocket connections the bill started climbing with every connection. We rebuilt it on Phoenix Channels and kept the same behaviour our clients already relied on.
Offline is the constraint that keeps disqualifying Elixir
Deliveries happen where there is no GSM at all. For those, the answer was a mobile app or a progressive web app collecting data offline and syncing later, not a Phoenix page.
None of these photos generalise
I put a picture of downtown San Francisco next to the Kenyan ones on purpose. Every scene in the talk is a specific place with specific problems, not a picture of a continent.
Watch the talk
Nailinda and Malcolm
Linda opened in 2018. She was learning to code and Malcolm was the friend she asked when she got stuck. He loved Python and he loved Elixir once he found he could reason about functional code. He also had sickle cell anaemia.
In Kisumu we had the drill. If Malcolm had an attack we called his father, his father called the physician, the physician came and the whole thing took ten or fifteen minutes. Kisumu is small. Everybody knew where to go.
Then Malcolm took a job in Nairobi. On his third trip he had an attack and the only hospital we were sure of was most of the way across the city. We sat in traffic for about an hour and a half. He was treated, we relaxed, we left to get a few things. Ten minutes later we got the call. He had died.
What we could not get past was that we had passed hospitals the whole way. We just did not know they were there. Nobody looks up the nearest hospital when they arrive somewhere, because nobody arrives expecting to get sick.
The name came before the product, on a whiteboard in the office. Nailinda, from the Swahili linda, to protect. People assume Linda named it after herself. She did not.
What it actually does
It asks three things: where you are, what the emergency is and a landmark. The landmark question is the one that took research. We collected the names people actually use, which turn out to be bars, parks and buildings rather than street addresses. Then it texts you the nearest hospitals and notifies those hospitals that you are coming.
The hospital has to answer. Either yes, come, we can handle this and we are expecting you, or no, we cannot, go elsewhere. That second answer is the point. It is worse to arrive, wait, then be sent on somewhere else.
At the time of the keynote we were in soft launch, with a handful of users and hospitals on the system, watching how many requests arrived and how errors surfaced. Asked outright whether Elixir had saved a life yet, Linda said no, not yet.
The next idea was to stop being the only front door. We started talking to a group campaigning against female genital mutilation, who were already handing out emergency numbers, about running their response on our infrastructure so they could concentrate on reaching people.
The farmer and 20 dollars
Different system, same country. A non-profit lets smallholder farmers register in advance for the next season and say what they intend to plant, so the long rains do not arrive before the seed does. Experts package the inputs. A package starts at around 20 US dollars and the farmer pays for it before collection.
Twenty dollars is a lot. Linda and I used to walk home from the office as a group to save the 30 cents that got us back in the morning, on days when lunch also cost 30 cents. So the organisation lets farmers pay in pieces, ten cents, thirty cents, a dollar, whenever they have it. Pay regularly enough to reach a threshold and you can collect without having finished paying.
Which is where the software comes in. When an old man turns up on delivery day and his package is not there, the question is which of several things went wrong. Payments that never landed. A seed that ran out. An officer who recorded a delivery that never happened. A qualification that was assessed wrongly. Somebody has to be able to answer that on the spot.
Traffic arrives in spikes, on deadlines
Registration closes and payment deadlines close. These are micro-payments and a large share of them land at 11:59pm on the last night, because that is how you keep your inputs.
Deliveries happen where there is no signal
Sometimes there is no GSM at the delivery site at all. Data is captured offline and synced back when the connection returns.
LiveView felt like a Polaroid
I asked Linda if she remembered instant cameras. Before them you took a photo and waited days for the negative to come back. Being able to validate everything as it happened is why the admin dashboard is LiveView.
The bottom ten
A third system, also Elixir, supports a programme running remedial classes in low-fee primary schools. The selection criterion is blunt: take the bottom ten of a class. Extra lessons run before the school day and during games time. Every session records what was taught, how the children did and who turned up.
At the end of the year that data answers questions the programme could not otherwise ask. Whether a tutor's own attendance tracks their students' attendance. Whether one tutor is doing better than another and on which measure. After a year or two, whether the whole programme is doing anything at all.
Freebird and a 15,000 dollar breakup
The commercial one. We were paying for a hosted real-time API so users could be told when they, or the people they were watching, went offline and came back. Past 100 websocket connections the bill started rising with the connection count. At 15,000 dollars we stopped and looked at each other.
Elixir had been promising us real time for a while. So we built the thing we were renting. Linda had asked, during mob sessions, where she would ever use Phoenix Channels. This is where.
We called it Freebird, after the Lynyrd Skynyrd song, because it was a breakup. The rule for the migration was that clients should not be able to tell. Same promise, same behaviour, no new complaints about being told they were offline while online. The bill after the switch was the cost of hosting our own servers.
It also handed us problems we did not have before and some of them we still had not solved when I gave the talk. One device going offline can produce seven to ten API calls, depending on how many nodes are up. That is the honest version. We cut the cost and dug ourselves a new hole.
Where things run
From the questions afterwards. Nailinda went to Heroku, where the documentation covers everything you need. Freebird runs on Kubernetes from Docker images. The farmer system deploys to AWS through GitLab CI/CD, packaged from a YAML file into containers.
The Elixir version at the time was 1.6 or 1.7, before releases were built in, so we used Distillery. That was a long week of working out what goes where.
We chose GitLab partly for cost. Back then it was the option that gave unlimited private repositories with CI/CD included, so we did not have to make a repository public to get a pipeline. A keynote the day before had put it well: get your CI/CD right and you have an extra employee. GitLab was ours.
The disclaimer
Near the end I showed a photograph of downtown San Francisco. One of the richest regions in the world and you still find that. Every image in the talk is a niche, a selected region and a selected problem that we happen to work on.
Then we showed where we live. Kisumu, on Lake Victoria, an active nightlife, good beer, one of the best sunsets I know of and people fishing on the largest freshwater lake in Africa. There are places in the city with poor internet. There is also enough internet that I streamed a conference keynote from the office.
Both are true. That was the whole point of putting them next to each other.
Think your organisation has outgrown its systems?
Let's figure out what is actually broken.