Showing posts with label nanopore. Show all posts
Showing posts with label nanopore. Show all posts

Saturday, April 22, 2017

Sunday Morning Insight: "No Need for the Map of a Cat, Mr Feynman" or The Long Game in Nanopore Sequencing.



About 5 weeks ago, we wondered how we could tell if the world was changing right before our eyes ?  well, this is happening, instance #2 just got more real:


Nanopore sequencing is a promising technique for genome sequencing due to its portability, ability to sequence long reads from single molecules, and to simultaneously assay DNA methylation. However until recently nanopore sequencing has been mainly applied to small genomes, due to the limited output attainable. We present nanopore sequencing and assembly of the GM12878 Utah/Ceph human reference genome generated using the Oxford Nanopore MinION and R9.4 version chemistry. We generated 91.2 Gb of sequence data (~30x theoretical coverage) from 39 flowcells. De novo assembly yielded a highly complete and contiguous assembly (NG50 ~3Mb). We observed considerable variability in homopolymeric tract resolution between different basecallers. The data permitted sensitive detection of both large structural variants and epigenetic modifications. Further we developed a new approach exploiting the long-read capability of this system and found that adding an additional 5x-coverage of "ultra-long" reads (read N50 of 99.7kb) more than doubled the assembly contiguity. Modelling the repeat structure of the human genome predicts extraordinarily contiguous assemblies may be possible using nanopore reads alone. Portable de novo sequencing of human genomes may be important for rapid point-of-care diagnosis of rare genetic diseases and cancer, and monitoring of cancer progression. The complete dataset including raw signal is available as an Amazon Web Services Open Dataset at: https://github.com/nanopore-wgs-consortium/NA12878.
Here is some context:

And previously on Nuit Blanche:
 
Credit: NASA, JPL




Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Wednesday, July 27, 2016

Streaming algorithms for identification of pathogens and antibiotic resistance potential from real-time MinION (TM) sequencing

About four years ago, I tried to predict the future for August 25, 2030. In order to do this, I first mentioned The Steamrollers i.e. technologies that were exponential in nature. Nanopore sequencing was one of them. In the second installment, I mentioned different algorithms that could help in making sense of the data generated by these steamrollers (Predicting the Future: Randomness and Parsimony). Streaming was one of them. It is really no surprise, if like in hyperspectral imaging or nanopore sequencing you are producing a lot of data, your interest switch from the modeling aspect of things to how can it be helpful now and how fast.  How does it change science ? well you just need to read the following article:

The main contribution of this article is to demonstrate that despite the higher error rate, it is possible to return clinical actionable information, including species and strain identification from as few as 500 reads. We achieved this by developing novel approaches that are less sensitive to base-calling errors and which use whatever subset of genome-wide information is observed up to a point in time, rather than a panel of pre-defined markers or genes. For example, the strain typing presence/absence approach relies only on being able to identify homology to genes and also allows for a level of incorrect gene annotation.



Streaming algorithms for identification of pathogens and antibiotic resistance potential from real-time MinIONTMsequencing by Minh Duc Cao, Devika Ganesamoorthy, Alysha G. Elliott, Huihui Zhang, Matthew A. Cooper and Lachlan J.M. Coin
The recently introduced Oxford Nanopore MinION platform generates DNA sequence data in real-time. This has great potential to shorten the sample-to-results time and is likely to have benefits such as rapid diagnosis of bacterial infection and identification of drug resistance. However, there are few tools available for streaming analysis of real-time sequencing data. Here, we present a framework for streaming analysis of MinION real-time sequence data, together with probabilistic streaming algorithms for species typing, strain typing and antibiotic resistance profile identification. Using four culture isolate samples, as well as a mixed-species sample, we demonstrate that bacterial species and strain information can be obtained within 30 min of sequencing and using about 500 reads, initial drug-resistance profiles within two hours, and complete resistance profiles within 10 h. While strain identification with multi-locus sequence typing required more than 15x coverage to generate confident assignments, our novel gene-presence typing could detect the presence of a known strain with 0.5x coverage. We also show that our pipeline can process over 100 times more data than the current throughput of the MinION on a desktop computer.






Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Monday, April 04, 2016

DeepNano: Deep Recurrent Neural Networks for Base Calling in MinION Nanopore Reads

In genomics, Base Callers are codes that figure out the G, T, A and Cs of DNA molecules after those have been cut into pieces in order to reassemble all that information back into one long string.

The nanopore technology promises to revolutionize genomics because, quite simply, it makes a formerly NP-hard problem of putting this information back together into anot so hard problem (P). This is why the nanopore technology is one of the steamrollers ( Predicting the Future: The Steamrollers ). Today, the data coming out of these sensors seem amenable to better classification thanks to Deep Learning thereby reducing its error rate and slowly putting on a par with other technologies. Woohoo ! Time for a Miller wave.

Let us note that these readings might have been looked at from the standpoint of a regular signal processing issue but people seem to eagerly try deep learning first. This is another example of the Great Convergence. Without further ado:




DeepNano: Deep Recurrent Neural Networks for Base Calling in MinION Nanopore Reads by Vladimír Boža, Broňa Brejová, Tomáš Vinař

Motivation: The MinION device by Oxford Nanopore is the first portable sequencing device. MinION is able to produce very long reads (reads over 100~kBp were reported), however it suffers from high sequencing error rate. In this paper, we show that the error rate can be reduced by improving the base calling process.
Results: We present the first open-source DNA base caller for the MinION sequencing platform by Oxford Nanopore. By employing carefully crafted recurrent neural networks, our tool improves the base calling accuracy compared to the default base caller supplied by the manufacturer. This advance may further enhance applicability of MinION for genome sequencing and various clinical applications.
Availability: DeepNano can be downloaded at this http URL
The website for the paper and code is here: http://compbio.fmph.uniba.sk/deepnano/
Let us also note that it looks like that even the nanopore sensor maker is also going that route.


 Previous blog entries:



Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Tuesday, October 27, 2015

Fast genome and metagenome distance estimation using MinHash / Machine learning for metagenomics: methods and tools

Awesome!  If you've read these blog entries [1,2,3] you know that genomics is now heading toward the analysis of large population of genomes aka metagenomics.

The first preprint was promised back in August and uses a Locality Sensitive Hashing to perform a comparison on a large population of genomes. The second preprint reviews several ML techniques for metagenomics and points to the earlier third preprint where another hashing technique is used.  That last approach used Vowpal Wabbit that is not a Locality Sensitive Hashing technique while MinHash used in the first preprint is. Without further ado:

Fast genome and metagenome distance estimation using MinHash by Brian D. Ondov, Todd J. Treangen, Adam B. Mallonee, Nicholas H. Bergman, Sergey Koren, Adam M. Phillippy

Machine learning for metagenomics: methods and tools by Hayssam Soueidan, Macha Nikolski

While genomics is the research field relative to the study of the genome of any organism, metagenomics is the term for the research that focuses on many genomes at the same time, as typical in some sections of environmental study. Metagenomics recognizes the need to develop computational methods that enable understanding the genetic composition and activities of communities of species so complex that they can only be sampled, never completely characterized.
Machine learning currently offers some of the most computationally efficient tools for building predictive models for classification of biological data. Various biological applications cover the entire spectrum of machine learning problems including supervised learning, unsupervised learning (or clustering), and model construction. Moreover, most of biological data -- and this is the case for metagenomics -- are both unbalanced and heterogeneous, thus meeting the current challenges of machine learning in the era of Big Data.
The goal of this revue is to examine the contribution of machine learning techniques for metagenomics, that is answer the question "to what extent does machine learning contribute to the study of microbial communities and environmental samples?" We will first briefly introduce the scientific fundamentals of machine learning. In the following sections we will illustrate how these techniques are helpful in answering questions of metagenomic data analysis. We will describe a certain number of methods and tools to this end, though we will not cover them exhaustively. Finally, we will speculate on the possible future directions of this research.
 

Large-scale Machine Learning for Metagenomics Sequence Classification by Kévin Vervier, Pierre Mahé, Maud Tournoud, Jean-Baptiste Veyrieras, Jean-Philippe Vert
Metagenomics characterizes the taxonomic diversity of microbial communities by sequencing DNA directly from an environmental sample. One of the main challenges in metagenomics data analysis is the binning step, where each sequenced read is assigned to a taxonomic clade. Due to the large volume of metagenomics datasets, binning methods need fast and accurate algorithms that can operate with reasonable computing requirements. While standard alignment-based methods provide state-of-the-art performance, compositional approaches that assign a taxonomic class to a DNA read based on the k-mers it contains have the potential to provide faster solutions. In this work, we investigate the potential of modern, large-scale machine learning implementations for taxonomic affectation of next-generation sequencing reads based on their k-mers profile. We show that machine learning-based compositional approaches benefit from increasing the number of fragments sampled from reference genome to tune their parameters, up to a coverage of about 10, and from increasing the k-mer size to about 12. Tuning these models involves training a machine learning model on about 10 8 samples in 10 7 dimensions, which is out of reach of standard soft-wares but can be done efficiently with modern implementations for large-scale machine learning. The resulting models are competitive in terms of accuracy with well-established alignment tools for problems involving a small to moderate number of candidate species, and for reasonable amounts of sequencing errors. We show, however, that compositional approaches are still limited in their ability to deal with problems involving a greater number of species, and more sensitive to sequencing errors. We finally confirm that compositional approach achieve faster prediction times, with a gain of 3 to 15 times with respect to the BWA-MEM short read mapper, depending on the number of candidate species and the level of sequencing noise.
 
 
 
Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Saturday, October 17, 2015

Sunday Morning Insight: Miller's Wave and Genomics

[Spoiler alert: if you have not seen the movie Interstellar then please do not read this post]




There is a scene in Interstellar that parallels what we are currently seeing in genomics. As Cooper and his crew lands on Miller's planet, the audience witnesses a watery landscape with what looks like far-away mountain ridges. Two minutes later, Cooper's spacecraft is nearly destroyed by a gigantic wave. 
 A minute it was a far away mountain ridge, the next it's a hundred story wave.. 
 
A week later, DNA.land has gathered 6000+ genomes.   

So the next thing we ought to ask ourselves besides producing even better sensors, is: Do we have algorithms that can scale for this sort of data stream on a daily basis ?

 
 
 
Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Friday, September 18, 2015

Entropy-Scaling Search of Massive Biological Data - implementation -

A continuation of the Compressive Genomics approach:
  


Entropy-Scaling Search of Massive Biological Data by Y. William Yu, Noah M. Daniels, David Christian Danko, Bonnie Berger



Highlights

  • •We describe entropy-scaling search for finding approximate matches in a database
  • •Search complexity is bounded in time and space by the entropy of the database
  • •We make tools that enable search of three largely intractable real-world databases
  • •The tools dramatically accelerate metagenomic, chemical, and protein structure search

Summary

Many datasets exhibit a well-defined structure that can be exploited to design faster search tools, but it is not always clear when such acceleration is possible. Here, we introduce a framework for similarity search based on characterizing a dataset’s entropy and fractal dimension. We prove that searching scales in time with metric entropy (number of covering hyperspheres), if the fractal dimension of the dataset is low, and scales in space with the sum of metric entropy and information-theoretic entropy (randomness of the data). Using these ideas, we present accelerated versions of standard tools, with no loss in specificity and little loss in sensitivity, for use in three domains—high-throughput drug screening (Ammolite, 150× speedup), metagenomics (MICA, 3.5× speedup of DIAMOND [3,700× BLASTX]), and protein structure search (esFragBag, 10× speedup of FragBag). Our framework can be used to achieve “‘compressive omics,” and the general theory can be readily applied to data science problems outside of biology (source code: http://gems.csail.mit.edu).

Obviously the implementation is here.
 
Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Thursday, September 11, 2014

A Detailed Overview of Nanopore Sequencing

Because the sensor technology is so important [0], I recently mentioned why long read technology was such a breakthrough (it's an information theoretic issue [5]). For the past two years, we have featured, here on Nuit Blanche, at least three makers of sequencing technology that will change the course of how we do many things. Those outfits are:
Nick Loman is one of the few people who has had access to Oxford Nanopore Technologies' long read technology through their early access program. To get a sense of how the technology will develop, he just organized a Hangout yesterday with Clive Brown of Oxford Nanopore Technologies. You really need to watch this video of the hangout to get a sense of the algorithms being envisionned and those already implemented for alignement purposes. In a future entry, I will talk specifically about Clive's presentation as I think there is potentially some additional information to be gained from their raw data in light of recent advances in compressive sensing and attendant algorithm development, stay tuned. Without further ado, here is the program of the video: 
  • Clive Brown, Nanopore sequencing 
  • Nick Loman, Early data from nanopore sequencing: bioinformatics opportunities and challenges 
  • Matt Loose, Streaming data solutions for nanopore 
  • Josh Quick, Nanopore sequencing in outbreaks 
  • Torsten Seemann, Awesome pipelines for microbial genomics 

and the video, enjoy !

 
Thank you Nick !
 
 
Join the CompressiveSensing subreddit or the Google+ Community and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Tuesday, February 25, 2014

Predicting the Future: The Upcoming Stephanie Events

You probably recall this entry on The Steamrollers, technologies that are improving much faster than Moore's law if not more. It turns out that at the Paris Machine Learning Meetup #5, Jean-Philippe Vert [1] provided some insight as to what happened in 2007-2008 on this curve:


Namely, the fact that the genomic pipelines went parallel. The curve has been going down ever since but one wonders if we will see another phase transition.

The answer is yes. 

Why ? from [3]
However, nearly all of these new techniques concomitantly decrease genome quality, primarily due to the inability of their relatively short read lengths to bridge certain genomic regions, e.g., those containing repeats. Fragmentation of predicted open reading frames (ORFs) is one possible consequence of this decreased quality.

it is one thing to go fast it is another to produce good results. Both of these issues of speed and accuracy due to short reads may well answered with nanopore technology and specifically the fact that entities like Quantum Biosystems Provides Raw Data Access to New Sequencing Technology and I am hearing through the grapevine that their data release program is successful. What are the consequences of the current capabilities ? Eric Schadt calls it the The Stephanie Event (see also [4]). How much faster are we going to have those events in the future if we have very accurate long read lengths and attendant algorithms [2, 5] ? I am betting more.

Sunday, September 22, 2013

Sunday Morning Insight: A conversation on Nanopore Sequencing and Signal Processing

From [8]

As some of you have noticed, nanopore sequencing is a subject that comes back often here. One of the latest instance is another Sunday Morning Insight on Thinking about a Compressive Genome Sequencer. All the entries on the subject can be found under the nanopore tag. Because I wanted to be more informed about the technology and see how compressive sensing could be inserted in it, I reached out to a few people to get some conversation going. 

The following is the result a long email exchange with someone also interested in this area of nanopore sequencing, in which we tried to clarify some of the potential issues for Nanopore signal analysis, based on the published results. The letter "I" is for my remarks and questions while "A" is for the person with whom I had this conversation and who wishes to remain anonymous. A big thank you to this person for the great conversation! 

I: Here are two or three things that bug me and I was wondering if you could provide some enlightenment.

From an outsider's point of view, the various reports I see about nanopore engineering seem contradictory. On the one hand, the voltage curves I see seem to have pretty low noise yet, I also see somewhere that the overall success for these techniques are only about 96% accurate( where one would expect 99.99% or better) These two facts do not fit. The only way I can reconcile them, so far, is as follows:
  • We see only the good traces with low noise, it's OK, it's PR. In reality, noise is actually much worse in general.
  • Voltage drift issues are not minimal in any sense of the word especially when reading long strands.
  • A false twin issue? As I previously mentioned in the blog (see http://nuit-blanche.blogspot.fr/2013/04/structural-information-in-nanopore.html ) a potential issue involving knots might be at play. If one were to assume that the distances between voltage steps changes are not uniform then it becomes difficult to make out the difference between, say, a nucleotide G that has a knot behind it (and which takes a little while to get through the nanopore) and two perfectly fine Gs following each other in the DNA strand. That way, there is pretty difficult classification issue.
  • Finally, for reasons that are still unknown, we see voltage step readings that can neither be classified as any of the A, G, T and C nucleotide, a situation that might be related to some of the issues mentioned above or a combination thereof or others (such as voltage change across the nanopore)
What is your feeling about this outsider's analysis or am I way off base? Is there is a simpler explanation for the low 96% accuracy?

By the way, the problems mentioned above are not insurmountable, it's just that we need to be more serious on taking a stab at them. One could, for instance, remove the slowly moving drift using an "analysis based dictionary learning" approach. There are actually other methods as well.


A: The first thing to note is that the paper you mention at:


Is an analysis of single bases free in solution, not strands of DNA. I don't believe any work showing a protein nanopore producing clear, single base resolution reads has been published (as an peer reviewed paper) to date.

The current measurements shown there do look quite clean. So, yes I can see why you'd expect a low error rate. There are probably a few sources of error you might want to consider:

  1. All reports I've seen show that the dwell time of molecule in a protein nanopore is exponentially distributed. This is why people in the ion channel literature people are mostly happy using HMMs, because it does appear to be a Markov process [1]. Given that the time is exponentially distributed, there's a high probability that you might not see a base, or see it for a short time. This is one possible source of error, in this case deletions.
  2. Most research talks about the need to control the motion of the DNA through the nanopore. This could be a significant source of error [3] [4] [5] [6].
  3. The diagram on your blog shows 4 bases in a current range of ~15pA, but in the literature no one has presented single base resolution on strands from a protein nanopore that I'm aware of. The work has shown “that several nucleotides contribute to the recorded signal” [6] [7].

I: I think the sentence in ref [6] makes it plain obvious as to why a low pass sensor might be interesting, from [6]

“thus far the reading of the bases from a DNA molecule in a nanopore has been hampered by the fast translocation speed of DNA together with the fact that several nucleotides contribute to the recorded signal”.

I: What is the actual purpose of these "motors" that control the motion of the strand? Is it that:

  1. with them the process is slowed down so that we have enough electrons per base (as you mentioned earlier if it goes faster we might be electron starved for the signal.) [we talked before about 1pA being 6 electronics per time interval at 1MHz sample rate].
  2. without them there would be no strand going through the pore?
  3. without them the nominal dwell time would be not very well defined?
  4. make sure that the knotty situation mentioned in the blog entry does not influence unduly the dwell time in the pore?
  5. any or all these explanations?
A: Possibly all of the above, depending on the system. Slowing down the strand (1) is probably the most significant contribution. If you think about the default case where there are no forces at play the DNA would be moving around under Brownian motion. There might be other local forces at play that make it move faster or slower. You might be able to control that motion with a "motor", but that might not always work very well.

In addition to this, some work shows that several nucleotides might contribute to the signal [7].

I: Please explain this last sentence. If it is what I think it is, it is very interesting.

A: So, to quote the wikipedia page on nanopore sequencing:

"In the early papers methods, a nucleotide needed to be repeated in a sequence about 100 times successively in order to produce a measurable characteristic change"

If you slow the strand down, or make other changes, you might still be faced with the problem that “several nucleotides contribute to the recorded signal” [6] [7].

I: Going back to the motor. With no motor, the strand can go up or down, and since you have access only to the current, you have really no idea which direction the strand is going.

A: This is correct.

I: Which brings me to a different type of question: is there other information gathered during those experiments that could be used to detect what direction the strand is taking? and at what speed? Are there some additional measurements made during those experiment?

A: No, I don't know of any additional measurements that could be made. I think some people have talked about using fluorescence to detect the motion of the strand through a pore but I don't know how far that works has gone.

If you slow the strand down, or make other changes, you might still be faced with the problem that “several nucleotides contribute to the recorded signal” [6] [7], i.e. signal does not come from a single position. Cherf et al.[7] is probably a good reference to look at for some example traces and information on this.

I: A-ah! So the measurement seem to be falling in this category of group measurements/group testing This is really what compressive sensing projects well into. I guess the main issues are:
  • the strand going through is a stochastic process with a poisson distribution (the motor makes that distribution to be more peaked)
  • we do not seem to have other measurements that could directly or indirectly provide some side information about the actual speed of the strand going through. In another blog entry (http://nuit-blanche.blogspot.com/2013/03/of-well-logging-and-nanopores.html) I made the parallel between nanopore and well logging/drilling issues. In the drilling issue, though, the probes have accelerometers on them so that a relatively simple kalman filter on top of the other information (akin to the current sensing in the nanopore) allows a much cleaner picture to emerge.
Is there anything else I am missing from that picture?

A: I think that's pretty accurate!

I: A final question on motors, do all nanopore systems need a motor?

A: Having a way of controlling the motion is desirable.

I: Are you telling me there are other ways of controlling the motion that do not require motors?

A: Speed can be controlled by various factors including:
  • Viscosity of the buffer
  • Applied voltage
  • Salt concentration
  • Temperature

They also suggest there that optical and magnetic tweezers and "DNA Transistors" could be used to control the actual motion so there are a bunch of options I think.


I: Ah! This is interesting, maybe I should get my hands on this book (Nanopores - Sensing and Fundamental Biological Interactions). That the voltage across the pore also change the dynamic is also worth investigating.

I: Thank you very much

Using the commenter's feedback, I went ahead and read this very well written 2011 review of the technology [8] (Nanopore sensors for nucleic acid analysis by Bala Murali Venkatesan and Rashid Bashir) with the following abstract:
Abstract: Nanopore analysis is an emerging technique that involves using a voltage to drive molecules through a nanoscale pore in a membrane between two electrolytes, and monitoring how the ionic current through the nanopore changes as single molecules pass through it. This approach allows charged polymers (including single-stranded DNA, double-stranded DNA and RNA) to be analysed with subnanometre resolution and without the need for labels or amplification. Recent advances suggest that nanopore-based sensors could be competitive with other third-generation DNA sequencing technologies, and may be able to rapidly and reliably sequence the human genome for under $1,000. In this article we review the use of nanopore technology in DNA sequencing, genetics and medical diagnostics
In the context of the discussion above, Here are some excerpts of the review of interest:
"... A structural drawback with α-haemolysin is that the cylindrical β-barrel can accommodate up to ~10 nucleotides at a time, all of which significantly modulate the pore current [25]: this dilutes the ionic signature of the single nucleotide in the 1.4 nm constriction, thus reducing the overall signal-to-noise ratio in sequencing applications…”
I: So in this instance, we have a group measurement and the signal-to-noise ratio definition is really about sensing a single nucleotide within a larger group.
“... Moreover, in experiments involving immobilized ssDNA, as few as three nucleotides within or near the constriction contributed to the pore current [27] compared with the ten or so nucleotides that modulate the current in native α-haemolysin [25]....”
I: Again the concept of group measurements.
“...Unidirectional transport of dsDNA through this channel (from amino-terminal entrance to carboxyl-terminal exit) was also observed [29], suggesting a natural valve mechanism in the channel that assists dsDNA packaging during bacteriophage phi29 virus maturation. The capabilities of this protein nanopore will become more apparent in years to come....”
I: The review highlights a possible mechanism to constrain the strand in only one direction.
“...The first reports of DNA sensing using solid-state nanopores emerged in early 2001 when Golovchenko and co-workers used a custom-built ion-beam sculpting tool with feedback control to make nanopores with well-defined sizes in thin SiN membranes [42]...”
I: This is one element I had not really understood, the possibility of having solid state nanopore (and potentially use Moore’s law).
“....Indeed, we observed that DNA translocation was slower in Al2O3 nanopores than in SiN nanopores with similar diameters, which was attributed to the strong electrostatic interactions between the positively charged Al2O3 surface and the negatively charged dsDNA [45]. Enhancing these interactions, either electrostatically or chemically, could reduce DNA velocities even more....”
I: or even control it **during** the analysis!
“...Translocation velocities were between about 10 and 100 nucleotides per microsecond, which is too fast for the electronic measurement of individual nucleotides…”
I: And this is where the idea of A2I comes out ( see Sunday Morning Insight: Thinking about a Compressive Genome Sequencer at http://nuit-blanche.blogspot.com/2013/08/sunday-morning-insight-thinking-about.html , use the architecture developed for these low pass sensors to get an idea of what passes through the solid state nanopore.
“...This result suggests that if the translocation speed could be reduced to roughly one nucleotide per millisecond, single-nucleotide detection should be possible, which could potentially lead to DNA sequencing with electronic readout…”
So this is, in my mind, a signal processing issue. Much discovery goes in developing hardware/,materials to slow down the phenomenon when one could probably look at it with current speeds and a different signal processing approach.
“....For example, is single-nucleotide resolution possible in the presence of thermodynamic fluctuations and electrical noise? And will the chemical and structural similarity of the purines (A and G) and the pyrimidines (C and T) inherently limit the identification of individual nucleotides using ionic current?...”
Looks like even the specialists are asking themselves good questions!
“.....SNPs and point mutations have been linked to a variety of mendelian diseases as well as more complex disease phenotypes [67]. In proof-of-principle experiments, SNPs have been detected using ~2-nm-diameter SiN nanopores [68]. Using the nanopore as a local force actuator, the binding energies of a DNA binding protein and its cognate sequence relative to a SNP sequence could be discriminated (Fig. 4b). This approach could be extended to screen mutations in the cognate sequences of various other DNA binding proteins, including transcription factors, nucleases and histones.....”

I: This is an interesting use of side information.
“.....Similarly, given the progress with solid-state nanopores, if the translocation velocity could be reduced to a single nucleotide (which is ~3Å long) per millisecond, and if nucleotides could be identified uniquely with an electronic signature (an area of intense research), it would be possible to sequence a molecule containing one million bases in less than 20 minutes....”
I: Again the reduction of speed to get “pure” signals
“....There have been preliminary reports on the use of embedded planar gate electrodes in nanopores [40] and nano-channels [81,82] to electrically modulate the ionic pore current, and the integration of single-walled carbon nanotubes for the translocation of ssDNA [83]. …”

I: It looks to me like one of the principal element of a low pass sensors descrived above from the A2I philosophy. Other mechanical changes or side information to the current device include:

“.....Recent experiments with scanning tunnelling microscopes suggest that it might be possible to identify nucleotides with electron tunnelling [89] (because the energy gaps between the highest occupied and lowest unoccupied molecular orbitals of A, C, G and T are unique [90]), and partially sequence DNA oligomers [91]....”
“.....Efforts to fabricate nanopore sensors that contain nanogap-based tunnelling detectors are currently underway [93,94], but thermal fluctuations and electrical noise present major challenges.....”
“.....Another challenge is the fact that tunnelling currents vary exponentially with both the width and the height of the barriers that electrons have to tunnel through, which in turn depends on the effective tunnel distance and on molecule orientation.....”
“....A four-point-probe measurement could therefore reveal significantly more information than the two-probe measurements attempted so far, but reliably fabricating such a four-probe structure with subnanometre precision will be a formidable challenge. It should also be noted that it is not necessary to uniquely identify all four bases for certain applications. Some researchers have used a binary conversion of nucleotide sequences (A or T = 0, and G or C = 1), to discover biomarkers and identify genomic alterations in short fragments of DNA and RNA [95,96]..."
From [8]

[2] James Clarke, Hai-Chen Wu, Lakmal Jayasinghe, Alpesh Patel, Stuart Reid, Hagan Bayley (2009). Continuous base identification for single-molecule nanopore DNA sequencing Nature Nanotechnology
[3] Controlled translocation of individual DNA molecules through protein nanopores with engineered molecular brakes, Marcela Rincon-Restrepo, Ellina Mikhailova, Hagan Bayley, and Giovanni Maglia.
[4] Nanopore Analysis of Nucleic Acids Bound to Exonucleases and Polymerases, David Deamer.
[6] “thus far the reading of the bases from a DNA molecule in a nanopore has been hampered by the fast translocation speed of DNA together with the fact that several nucleotides contribute to the recorded signal”. DNA sequencing with nanopores, Grégory F Schneider & Cees Dekker, Nature Biotechnology. http://ceesdekkerlab.tudelft.nl/wp-content/uploads/Nature.pdf
[7] Automated forward and reverse ratcheting of DNA in a nanopore at 5-Å precision, Cherf et al. Nat. Biotechnol. 30, 344–348 (2012).
[8] Nanopore sensors for nucleic acid analysis by Bala Murali Venkatesan and Rashid Bashir, Nature Nanotechnology, 6, 615–624 (2011). Published online 18 September 2011 also at: http://libna.mntl.illinois.edu/pdf/publications/127_venkatesan.pdf





Join the CompressiveSensing subreddit or the Google+ Community and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Saturday, September 07, 2013

Saturday Morning Videos: Big Data Boot Camp, Alyosha Molnar on Angle-Sensitive Pixel-Based 3D Camera, Nanopore processes, Aggie Robot

A comment was left on the Around the blogs in 78 hours blog entry about the The "Big Data boot camp" at the Simons Institute

All videos recordings of the bootcamp lectures so far are currently available at
they will soon been added to the official workshop website as well, which is
Thank you anonymous commenter. I'll keep you posted when they are listed or you can leave a comment here if you spot them there.

On his blog, Vladimir mentioned a video by Alyosha Molnar on Angle-Sensitive Pixel-Based 3D Camera
In this Youtube video Alyosha Molnar, Assistant Professor at the School of Electrical and Computer Engineering, Cornell University presents his ideas on angle-sensitive pixels for 3D imaging. The presentation appears to be based on the previously published slides here.


We mentioned some of this work before

I mentioned nanopores and how they work before. Here is a video that explains more vividly different scenarios (courtesy of the Oxford Nanopore Vimeo feed):

 
Run Until - DNA sequencing informatics on the GridION and MinION systems from Oxford Nanopore on Vimeo.
 
Protein Detection from Oxford Nanopore on Vimeo.
 
MiRNA Detection from Oxford Nanopore on Vimeo.

Finally, of interest an agriculture robot:


Printfriendly