The client and the context
Ochre Bio is a biotech founded in Oxford in 2019. It develops RNA therapies for liver diseases. Its labs produce a lot of genomic data, and a team of computational biologists analyzes it to identify the genes a treatment could target.
To browse these results, the team had started an R Shiny application. A scientist enters the name of a gene and gets a report showing in which liver cells the gene is expressed, which clinical observations it is associated with and what happens when it is silenced.
We joined the project in 2022, for the interface. Jörn Schmiedel, lead computational scientist at Ochre Bio, described the goal in a testimonial.
We wanted to build an interface that aggregates all the data and is able to display it to all scientists in the company but also to external stakeholders, investors, and really the public.
Our work then extended from the interface to the back end, then to the administration of the AWS account, and the application took the name OBELiX.
The problem
In late 2023, Ochre Bio planned to open OBELiX the following year to partner pharmaceutical companies. Their researchers would log in to the application to consult Ochre Bio’s data.
This new audience changed the requirements. Each partner had to see only the data it was entitled to. The application had to respond quickly, including when several teams were connected at the same time. The infrastructure hosting it had to meet the security rules of a large pharmaceutical company.
What we did
Charts computed in advance. So that the application responds quickly to several teams at once, we moved the chart computation out of it. Automated jobs produce every chart for every gene and save it to S3, a storage service. When a user opens a report, the application only has to display charts that are already there.
The trade-off is that everything has to be produced before opening. The database has 63,187 genes, which made about 630,000 charts with five types per gene and two user groups at the start. We therefore split the computation so that each chart can be redone on its own, in a few seconds, and a job that fails is restarted automatically.
Access rights set per group. Each user belongs to a group, either Ochre Bio’s or a partner’s. A table states, for each group, the sections of the report it is entitled to. The interface uses it to hide the unauthorized sections, and the computation jobs use it to skip their charts. A chart that a partner is not entitled to therefore does not exist for its group. The data too is stored in one folder per group.
Hosting rebuilt on AWS. We redid the infrastructure so that it meets the security rules of a large pharmaceutical company. The test environment and production live in two separate AWS accounts, and a change is verified in the first before reaching the second. The whole infrastructure is described in code, with CloudFormation, which makes it possible to recreate it identically and to review every change. It is deployed in two regions, so that the application stays available if one of them goes down. Partners log in with two-factor authentication, and Ochre Bio’s cloud provider monitors the application around the clock, with alarms and procedures prepared with that provider.
The challenge: twelve hours for each data update
Ochre Bio’s scientists regularly deliver a new version of the OBELiX data, as files. At first, each version had to be loaded into a database before the charts were recomputed. That loading took at least twelve hours, and an error spotted afterwards in the data meant starting over.
This delay became a problem a few days before opening to the first partner, when the data still had to be updated, and the charts with it. At twelve hours per attempt, there was room for very few corrections, on data the partner would consult from day one.
So we removed the loading step. The computation jobs read the files delivered by the scientists directly, in Parquet format, with DuckDB. The change was working the day after it was proposed. It was put to use a few days later, when two errors were found in the data: each was corrected and verified in a few minutes. Reading the files directly has remained how OBELiX works.
The results
OBELiX went from an internal tool to a platform that partner companies consult. It has been in production since July 2024, and Ochre Bio opened it to its first partner’s teams that same month. Six months later, a second partner had its own version, with its own home page and access rights.
Ochre Bio then took over maintenance in-house, with a developer whose onboarding we supported. We left in the repository the documentation and the procedures for routine operations, such as setting a group’s rights or rerunning the computation after a data update, and our engagement ended three months after the developer arrived.
More than a year later, the Ochre Bio team confirmed to us that the application was still running, without major changes, with far more data than when we left.
Do you have an internal Shiny application to open to clients or partners? Get in touch.