Atlas OS · Analysis and delivery
The complete chain
Between raw data and a delivered result there used to be manual work. Since this month that stretch runs as a machine – with the run as an object in the database, a checksum for every file, and a release that a person still grants.
On 22 August 2026 at 11:06, the last open stage of a whole-genome run reported the state “finished”. Behind it lie four days, 103.6 gigabytes of result data and not one manual touch of the data. Everything that happened during those four days was in our database before a person had read it: which tool at which version, on which compute node, from when to when, with which checksum going in and which coming out.
9
stages in one run, each logged separately
4
days from raw data to the finished delivery bundle
103.6 GB
of result data delivered, with a checksum per file
6
delivery lines with a checksum, each released on its own
0
manual interventions in the data between start and delivery bundle
Nine stages · four days · no manual touch
01
The gap sat in the middle
The front of the chain has been in place for months and has been written up: an order becomes a quotation, the quotation becomes an order confirmation, the confirmation becomes sample logistics with shipment tracking, the sample becomes an arrival in the lab. The back is in place just as much: portal, documents, invoice, release. What lay in between is run as manual work across our industry to this day. Raw data becomes alignments, alignments become variants, variants become an annotated and filtered selection, that becomes a report, and that becomes a delivery. Every one of those handovers was a person carrying a file from one place to another – and having to explain afterwards, from memory, where it came from.
- Quote
- Confirmed
- In transit
- At the lab
- Data delivered
- Completed
Order phase: Data delivered – Results released and available to download.
That middle is closed. For the first time the chain runs from the order through to a downloadable result without anyone pushing it along by hand at some point.
02
The run is an object now, not a story
Until now a compute run was something that happened on a machine and existed afterwards only as a folder. Today it is a record: pipeline and version, compute node, period, key figures, plus every single stage with its own tool, its own version, its own time window and its own state. The run hangs off the order, and every result file hangs off the run.
That sounds like bookkeeping. It is the difference between a result you have to take on trust and a result that brings its own provenance with it. Anyone asking in two years' time how a particular variant ended up in a particular table gets a row as the answer, not a recollection.
QN-2026-9xxx_P2026-3xx
Completed- Pipeline
- nf-core/sarek 3.8.1
- Period
- 18 Aug 2026, 19:54 – 22 Aug 2026, 11:06
- Compute node
- atlas-p0
- Stages
- 9
Run completed without intervention. All nine stages logged with tool, version, time window and checksum.
| Stage | Tool and version | Period | State |
|---|---|---|---|
| Alignment | bwa-mem 0.7.18 | 18 Aug 20:48 – 19 Aug 11:43 | Completed |
| Duplicate marking | gatk4 MarkDuplicates 4.6.1.0 | 19 Aug 11:44 – 13:36 | Completed |
| Recalibration | gatk4 BaseRecalibrator/ApplyBQSR | 19 Aug 13:38 – 16:21 | Completed |
| Variant calling | gatk4 HaplotypeCaller 4.6.1.0 | 19 Aug 16:21 – 22:52 | Completed |
| Annotation | Ensembl VEP 115 | 20 Aug 11:22 – 13:30 | Completed |
| Candidate filter | kandidaten.sh | 20 Aug 13:30 – 13:34 | Completed |
| Additional analysis | samtools/mosdepth, Ensembl REST | 20 Aug 14:08 – 21 Aug 15:29 | Completed |
| Splice scoring | Pangolin 1.0.2 | 20 Aug 22:34 – 22 Aug 11:06 | Completed |
| Structural variants | Manta 1.6.0 | 21 Aug 00:50 – 01:13 | Completed |
03
The machine navigates, the person decides
Every stage is selected, parameterised, started, monitored and evaluated without anyone setting it off. More interesting than running through is stopping. In one of the runs, a coverage weakness in a difficult genomic region could have stayed invisible under averaged window values. The analysis noticed the discrepancy itself, repeated the calculation at interval level, put both results side by side and wrote the rule that followed from it into the log. In another case a scanned document could not be read; instead of filling in the missing details, the run reported what it had been unable to determine and took the route via image recognition.
That is the actual point. A machine that only runs through is a risk. A machine that knows where its statement ends is a tool.
At the end of every chain there is therefore no automaton but an order of steps. Bioinformatics checks the run against its key figures. The technical and scientific team checks the report against the data. Where the case calls for it, medical sign-off goes on top. What has disappeared are the manual steps. What has remained are the decisions – and those are meant to remain.
How much falls away is best shown by a run itself. The annotation before it writes one row per transcript and one per gene for every variant – 45,042,328 and 6,632,525 rows. The cascade below counts variants, not rows, and therefore starts again from the beginning.
4,981,959variants from variant calling
261,763after frequency threshold
933after consequence class
795candidates in the delivered table
04
Nothing leaves the building that has not been released
Finished files are not sent, they are fed in. Every line carries a name, a size and a checksum, and it sits withheld at first. Only a switch on each line releases it to the customer, and that switch writes into the order history who released what and when. An internal working draft can therefore sit in the same order as the customer document without ever becoming visible.
For large data a rule applies that we took from providers who move billions of objects a day: the portal authorises, it does not transport. No shared folder password, no link that travels on. The address is created at the moment of the click, is valid for minutes or hours depending on file size, lets an interrupted transfer resume, and leaves a row behind on every retrieval. For the customer, six kilobytes and a hundred gigabytes therefore look the same.
Analysis report on the genome analysis, 11 pages
QN-2026-9xxx_auswertungsbericht.pdf
ReleasedCandidate variants with frequency, consequence and database comparison
QN-2026-9xxx_kandidatenvarianten.tsv.gz
ReleasedQuality report of the sequencing and analysis run
multiqc_report.html
ReleasedSupporting files and individual analyses, 45 files
QN-2026-9xxx_belege.zip
ReleasedInternal working draft for the lab team, not for distribution
QN-2026-9xxx_arbeitsfassung.docx
Withheld
Raw data and analysis files of the delivery, ten files, 103.6 GB
05
Why this holds for every question
The chain does not ask what biology stood at the beginning. Raw data comes in at the front, defined deliverables fall out at the back, and in between lies an interchangeable stretch. Whether a genome, an exome, a trio, a microbiome, an expression profile or a protein panel from the partner network stands at the start changes the tools in the middle. It changes nothing about what stands around them: run object, stages, manifest, checksums, withholding, release, log.
That is exactly why the effort is one-off and the benefit recurring. A new question costs a pipeline, not a new organisation.
06
The evidence
One order has run through this chain in full, with three further runs over the same stretch. The fourth case comes from an ongoing validation project and is not a genome but an exome – enriched rather than sequenced in full, and from a different type of sequencing instrument than the three before it. The same stretch, a different data source, a different method, without a line of special handling.
Four cases are no proof of throughput. They are something else, and at this moment that is worth more: evidence that the architecture holds. What runs through four times without special handling also runs four hundred times – the difference is then compute time, not organisation.
What this post leaves out, it leaves out on purpose. No customer is named, no finding appears, and where a collaboration is still under way, we wait for the people involved before we show anything. The people whose data we process decide when it is spoken about.
07
The rest of the way
Two things are still manual work today, and both are on the list for the coming weeks. Someone starts the run, instead of arriving raw data setting it off by itself. And someone looks at the finished document before it goes out, because a report is more than the text in it. The laboratory side has already gone the opposite way: the instruments have long been programmed rather than operated, and the handling of the samples will follow before long.
The ambition is immodest and worth writing down for exactly that reason. Nine hundred and ninety-nine runs in a thousand should pass through without intervention. The thousandth should stand out because a stage says so itself – not because someone happened to look.
Our measure is not how fast we produce files. It is the time from a biological question to a confirmed result, and the number of hands that have to reach in along the way.
Order, run and delivery lines sit in the portal at app.atlasbiolabs.com. How the front of the chain comes about – from the question to the binding quotation – is written up here.