The client and the context
Viking Conseil is a French company that designs electricity consumption forecasting models. It grew out of a PhD thesis on adaptive forecasting, defended in 2022, and it now sells its forecasts to electricity suppliers and grid operators.
These clients go through an API. Each day, it receives a client’s latest consumption measurements and updates that client’s model. It then receives the weather data for the coming days and returns the forecast. The clients have their own tools and use only this API.
We built it in 2023, along with a Shiny application. Since then, we have maintained it, with a report sent every month, and we develop it as Viking Conseil’s business grows. Its data is stored in DuckDB, a database that fits in a single file, and its server has been upgraded three times as models and data were added.
The problem
In late 2025, Viking Conseil onboarded several clients. All of them need their forecast in the morning, and calls are concentrated between 9 and 11 a.m.
The API is built on plumber, the package that exposes R code as an API, and it runs in a single process. It handles one request at a time, and the following ones wait their turn. The first timeout errors appeared during a test in which several models had to be updated in a row.
The founder of Viking Conseil then asked us how to make the API able to process several requests in parallel. At that point, we had no measurement of response time in production.
What we did
We proposed three independent steps: measure, fix what the measurements showed, then redesign the architecture if the need remained. Viking Conseil approved the first two.
Measure before rebuilding. A redesign would have run several copies of the API in parallel. Before committing to that work, we wanted to know where the time of a request went. So we connected the API to our monitoring infrastructure, where OpenTelemetry records the duration of each step of a request and Grafana displays it in a dashboard that the Viking Conseil team consults just as we do.
After a week of measurements, the breakdown of a forecast request showed where to look. It took 1 min 46 s, of which 18 seconds were computation and 1 min 28 s were spent reading the database. The API opened and closed the database for each of its nine reads, and each round took about 9 seconds. This behavior dated from the time when the Shiny application shared the database with the API.
Fix what the measurements point to. The fix came down to two changes, deployed at an off-peak hour. The API now keeps a single connection open to the database, and it loads all of a client’s information in one go. For Viking Conseil’s clients, the API looks the same as before, with the same addresses and the same access key, and they had nothing to change on their side.
The challenge: a side effect the next day
The day after the deployment, the Viking Conseil team noticed that the database file, which it copies regularly for its analyses, did not contain the morning’s operations. They only appeared in it in the early afternoon.
This was a consequence of the previous day’s change. With a connection open continuously, the database first records writes in a log, then transfers them in batches to the main file. No data was missing, and the morning’s forecasts had been sent to the clients. We explained the mechanism to the team the same day, and showed how to get a complete copy.
The results
Two weeks after the deployment, we compared the dashboard measurements before and after.
| Before | After | |
|---|---|---|
| A forecast | 2 min 05 s | 6 s |
| A model update | 3 min 37 s | 55 s |
| Wait before a request is processed | 26 s | 1.5 s |
| Time the API spends processing requests | 30 h per week | 4.5 h per week |
At peak hours, the API was busy up to 80% of the time, and those peaks have disappeared. The founder of Viking Conseil confirmed to us that the improvement was visible on the company’s side.
Since then, Viking Conseil has continued to onboard clients and deploy heavier models. The time of a request is now the time it takes to compute the models.
The dashboard and the monthly maintenance report track this increase in load. The third step, the architecture redesign, is already specified, and it will be launched when the measurements call for it.
Does your API or R application slow down as usage grows? Get in touch.