DNA sequencing has transformed healthcare by helping scientists identify the genetic causes of many diseases and enabling more targeted patient treatments. While whole genome sequencing is now considered an efficient, standard lab technology, I recall from my Master’s and PhD studies around 2010 that it often failed to capture large portions of the DNA, and the associated bioinformatic analysis was slow and expensive. Looking back, it has been remarkable to witness how short-read sequencing has helped drive the rise of multi-omics, from transcriptomics, and translatomics to epigenomics, reshaping our understanding of molecular processes and advancing biologics manufacturing at Lonza.
The era of long-read sequencing at Lonza
In recent years, long-read sequencing has become an important complement to short-read sequencing, offering new possibilities for genetic characterization. Its ability to generate substantially longer sequence reads enables a more comprehensive and accurate analysis of complex genomic regions.
In particular, long-read sequencing is well suited to detecting large structural genomic variants, such as insertions, deletions, inversions, and translocations, which can be difficult to resolve using conventional short-read approaches.
To put this into perspective, during my PhD I was excited to work on one of the few studies using paired-end reads of 101 base pairs (bp). Today, long-read sequencing using the Oxford Nanopore platform typically yields a median read length of around 30,000 bp from our CHO (Chinese Hamster Ovary) cell lines at Lonza, with some reads extending beyond one million bp.
Enhancing biologics development through long-read sequencing
At Lonza, we see long-read sequencing as a valuable way to better understand therapeutic antibody production in CHO cell lines. The journey begins by introducing vectors carrying the genes or “transgenes”, that encode the therapeutic molecule into CHO cells. These cells are then used to generate stable clones that can consistently produce the biologic drug. However, the number and genomic locations of these transgene insertions can vary, particularly when transposon-mediated integration is used, as in Lonza’s GS Gene Expression System® platform.
Characterizing transgene copy number, genomic integration sites, and the genetic integrity of each integrated sequence provides critical information for the ongoing optimization of our expression system. This can also support our customers’ Biologics License Application (BLA) submissions.
Historically, specialized short-read sequencing approaches, such as Targeted Locus Amplification (TLA-seq), have been used to determine the number and genomic locations of transgene insertions. Long-read sequencing now offers an opportunity to build on these approaches, with several technical and practical advantages. In particular, long reads can span entire transgene insertion sites and their surrounding genomic regions, enabling more comprehensive characterization of individual integration events. They can also, provide additional genome-wide information, including DNA methylation profiles, which may offer further insight into gene activity.
Demonstrating accuracy and reliability in our updated genetic characterization approach
As with any new technology, there are challenges that must be addressed through rigorous scientific evaluation. To achieve robust and reproducible analysis, we developed a Nextflow-based bioinformatic pipeline that processes raw data generated by nanopore sequencers, performs quality control, and provides detailed insights into transgene insertions within the CHO genome.
We then evaluated this pipeline using simulated sequencing reads that reflect antibody‑expressing CHO clone genomes across different molecule formats. In these simulations we varied the sequencing depth as well as the number and genomic location of transgene insertions. This approach allowed us to determine minimal data requirements and confirm that the bioinformatic pipeline outputs aligned with expected results from the simulations. In addition, by comparing short‑read TLA‑seq and long‑read nanopore data generated from the same cell lines, we observed a strong correlation in insertion numbers, further supporting the reliability of the long‑read approach.
For me, this work shows how advances in sequencing technology can be translated into practical tools for biomanufacturing and regulatory support. The roll-out of our end-to-end genetic characterization workflow highlights our commitment to innovation and strengthens Lonza’s position as a trusted partner in biomanufacturing.