Showing posts with label CompressiveSensingWhatIsItGoodFor. Show all posts
Showing posts with label CompressiveSensingWhatIsItGoodFor. Show all posts

Wednesday, December 17, 2014

More Oil and Compressive Sensing Connections



Felix Herrmann followed through to yesterday's "Another Donoho-Tao Moment ?" which itself was a followup to Sunday Morning Insight: The Stuff of Discovery and Hamming's time: Scientific Discovery Enabled by Compressive Sensing and related fields. Here what he has to say:

Hi Igor,
I agree with your assessment that the application of CS to seismic data acquisition may not qualify as a major scientific discovery but it will almost certainly lead to major scientific breakthroughs in crustal and mantle seismology as soon as my colleagues in those areas get wind of this innovation. Having said this it may be worth mentioning that randomized subsampling techniques are also having a major impact on carrying out large scale inversions in exploration seismology. For instance, a major seismic contractor company has been able to render full-waveform inversion (the seismic term for inverse problems involving PDE solves) into a commercially viable service by virtue of the fact that we were able to reduce the computational costs of these methods 5—7 fold using batching techniques. These references
were instrumental in motivating the oil & gas industry into adapting this batching technology into their business.
An area that is closer in spirit to Compressive Sensing is seismic imaging. Contrary to full-waveform inversion, seismic imaging entails the inversion of extremely large-scale tall systems of equations that can be made computationally viable using randomized sampling techniques in combination with sparsity promotion. While heuristic in nature, this technique
is able to approximately invert tall systems at the cost of roughly one pass through the data (read only one application of the full adjoint). It is very interesting to see that this heuristic technique has, at least at the conceptual level, connections with approximate message passing (see section 6.5 of “Graphical Models Concepts in Compressed Sensing”) and recent work on Kaczmarz: “A sparse Kaczmarz solver and a linearized Bregman method for online compressed sensing”.
While these developments may by themselves not qualify as major “scientific breakthroughs”, it is clear that these developments are having a major impact on the field of exploration seismology where inversions as opposed to applications of “single adjoints” are now possible that were until very recently computationally unfeasible. For further information on the application of Compressive Sensing to seismic data acquisition and wave-equation based inversion, I would like to point your readers to our website, mind map, and a recent overview article
Thanks again for your efforts promoting Compressive Sensing to the wider research community.
Kind regards,
Felix J. Herrmann
UBC—Seismic Laboratory for Imaging and Modelling
Thanks Felix !
 
Join the CompressiveSensing subreddit or the Google+ Community and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Tuesday, December 16, 2014

Another Donoho-Tao Moment ?

We always hear about the need to justify mathematics as regards to applied work. Remember this "Donoho-Tao moment"? (from this 2008 newletter)
[Mark Green's] favorite moment of the program came when the NSF panel – which was simultaneously conducting its first site visit – asked: Is this interdisciplinary work? Participant David Donoho (statistics, Stanford), another major contributor to the genesis of compressed sensing, reportedly exclaimed, “You’ve got Terry Tao talking to geoscientists, what do you want?”

Here is one that looks like another instance of it: Chuck Mosher, a Geoscience Fellow at ConocoPhillips mentioned the following in the LinkedIn thread related to scientific discovery using compressive sensing. While this is not a scientific discovery per se, it does fall in the discovery area and technology improvement: 

Igor,

I've recently started following Nuit Blanche, and find your site to be a great reference for compressive sensing topics.

I believe I have a possible example of CS leading to "discoveries" in my field: seismic data acquisition, processing, and imaging. I have worked with Rich Baraniuk at Rice University on many topics over the past 20 years or so. Rich recognized that seismic data acquisition might be a good candidate for compressive sensing innovation, and began coaching me in this direction starting around 2008 when CS was gathering steam. In 2010, I started a formal research project at my company (ConocoPhillips) to investigate applications of CS to petroleum exploration and production. In 2012, we acquired our first field trials using CS designs for seismic data acquistion, and finally in 2014 we deployed our first full scale CS based system for seismic data acquistion. We call our framework Compressive Seismic Imaging, or CSI. Gotta have a good acronym ;-)

Seismic data acquisition for a single geologic prospect can cost anywhere between $10-$100MM USD, and involves the use of the largest moving objects ever created by man (6 x 18 km sensor arrays with upwards of 64,000 channels). Compressive Sensing allows us to increase what we call "acquisition efficiency" by a factor of 2 or more in each coordinate direction that is employed. For seismic data acquisition, we have 4 coordinate directions (source x, source y , receiver x, receiver y), so efficiencies on the order of 2**4 or 16x are possible. In addition, CS can be used to enable the use of multiple simultaneous sources, providing another factor of 4 or so. Just as in parallel computing, a significant portion of these projects is "serial", so the efficiencies might only have a 2-10x impact on cost. You do the math - impacting cost with a factor of 2 on a $100 million dollar project will get your attention.

Capital programs for seismic data acquistion exceed $1 billion dollars for many of the large oil companies, so there is a lot of upside for CS in our business. We have only started to get this technology into use, but we expect that in a few years the seismic acquistion business will fully embrace CSI.

Here is a link to a feature article on our CSI program in a popular news magazine for the exploration geophysics industry, "The Leading Edge", published by the Society of Exploration Geophysicists:

http://www.tleonline.org/theleadingedge/april_2014?pg=28#pg28

Best regards,
Chuck Mosher
Geoscience Fellow, ConocoPhillips  
 
So this is the applied part. What about the pure math connection ? Dustin Mixon recently wrote about Alexander Grothendieck's influence on his work: 
(I say that his is not my field of study, and yet I have still seen his influence. For example, the Grothendieck inequality is a beautiful result in functional analysis that he proved as a graduate student before changing fields, and it has since found applications in hardness of approximation. Also, his development of Grothendieck groups provided a starting point for K-theory, which is the source of the best known lower bound for the 4M-4 conjecture.)


see also Phase Retrieval from masked Fourier transforms. In summary, you've got one of Grothendieck's theory providing bounds on the number of sensors and enabling a billion dollar industry, what do you want ?


Join the CompressiveSensing subreddit or the Google+ Community and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Sunday, December 14, 2014

Sunday Morning Insight: The Stuff of Discovery



In machine learning, they call the connundrum "exploitation versus exploration", in other circles we talk about "improving stuff versus discovery". There are many ways discovery can be defined. For instance in Crossing into P territory, we noted that a new kind of sensor (a genome sequencer) could enable experimentations in polynomial time thereby clearly expanding the possibilities to do real discovery. While both the PacBio and Oxford Nanopore technologies have just been made available before the summer, they are already changing the nature of discovery in that field [1]. Before, people would wonder how one could put pieces of DNA together, now that this complexity is mostly gone the new norm is now becoming: since genomes can be assembled easily, what sort of discovery can be done with a collection of genomes.
As I've said before, we have had this exact same explosion unraveling in compressive sensing ten years ago. What happened since ? Many polynomial time algorithms were developed with the emphasis of being faster than the previous ones and soon enough more complex data structures began to be exploited by the algorithms. There is really no reason to believe why this should not happen in genome sequencing: We are going to have many algorithms that do alignement using either PacBio or Oxford Nanopore technologies.

But in compressive sensing, something else happened: We began to discover a few things. This past Friday ( Hamming's time: Scientific Discovery Enabled by Compressive Sensing and related fields ) I provided two candidates. One candidate used the structure of the problem and its limitation to make predictions, while the other used the new paradigm to show how to nix a theory. Here are two other: An inference based on sparsity priors [2] (as featured in Catching an aha moment with compressive sensing ) and another one [3] about finding a needle in a haystack in an exponential families of solutions featured in ( Cluster expansion made easy with Bayesian compressive sensing ).

In actuality, those four examples fall into two categories: One category is where one uses the new prior as a way to find that needle in an exponential haystack while the other category uses empirical complexity bounds to reduce the phase space of what is feasible.

Either exploit the newfound capability or reduce the exploration horizon: different sides of the same discovery coin. Both are important.


References:
[1] Genomic sequencing: Recent tweets, papers, blog posts and attendant comment on that blog post:
Widespread polycistronic transcripts in mushroom-forming fungi revealed by single-molecule long-read mRNA sequencing by Sean Gordon, Elizabeth Tseng, Asaf Salamov, Jiwei Zhang, Xiandong Meng, Zhiying Zhao, Dongwan Don Kang, Jason Underwood, Igor V Grigoriev, Melania Figueroa, Jonathan S Schilling, Feng Chen, Zhong Wang


MinION nanopore sequencing identifies the position and structure of a bacterial antibiotic resistance island by Philip M Ashton, Satheesh Nair, Tim Dallman, Salvatore Rubino, Wolfgang Rabsch, Solomon Mwaigwisya, John Wain & Justin O'Grady
And a comment following this article: USB-sized DNA sequencer is error-prone, but still useful

Hi! Thanks for the write up! Getting spoken about on Ars Technica is definitely crossed something off my bucket list (I'm first author on the paper discussed).

I would just like to say a few things about the MinION/Oxford Nanopore:

1) While the error rate we observed is high compared to e.g. Illumina, it is comparable to PacBio (the main high throughput, long read tech).

2) You say 'the great promise of nanopore sequencing has been very difficult to match in practice'. However, I don't really think that is true. What ONT have done is amazing!

2a) First of all, the form factor is revolutionary. I'm not sure what your definition of a USB product is, but mine would be 'something where the only connection is a USB connection'. The MinION meets this.

2b) Rather than limiting the MinION device to a small number of elite institutes, they sent it to hundreds of 'normal' people. This is a brave move that speaks to the confidence they have in their technology. We had a positive experience with it, some people probably less so, others more so. This approach to letting everyone have a crack is surely one to applaud?

2c) The technology is just fantastic - single molecule sequencing using a biological pore! Think about how hard that must be to engineer! In a way that can be shipped to and used by hundreds of non-specialists! I should say that Illumina and PacBio also have awesome devices/technologies, but this one is newer ;-)

3) A slight technical issue, but the short reads weren't used to correct the long reads. The long reads were used to join contigs made using the short reads.

4) Another slight technical issue, the Illumina technology with bias is specifically the Nextera protocol. This has been known since this technology was developed by Jay Shendure's lab.

Thanks again for writing us up! 

 
[2] Direct inference of protein–DNA interactions using compressed sensing methods by Mohammed AlQuraishi, and Harley H. McAdams (featured in Catching an aha moment with compressive sensing )

[3] Lance J. Nelson*, Vidvuds Ozolins, C. Shane Reese, Fei Zhou, Gus L. W. Hart, "Cluster expansion made easy with Bayesian compressive sensing," Phys. Rev. B 88, 155105 (Oct. 2013). [pdf] and Lance J. Nelson*, Gus L. W. Hart, Fei Zhou, and Vidvuds Ozolins, "Compressive sensing as a paradigm for building physics models," Phys. Rev. B 87 035125 (2013). [pdf] featured in ( Cluster expansion made easy with Bayesian compressive sensing )
 
 
Join the CompressiveSensing subreddit or the Google+ Community and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Friday, December 12, 2014

Hamming's time: Scientific Discovery Enabled by Compressive Sensing and related fields



I was recently asked an interesting yet challenging question. I have two or three answers and also feel the question is somewhat unfair but it picked my interest. The question was: 
Has compressive sensing helped in making discoveries in the realm of general science (as opposed to computer science, signal processing, better solvers, etc...) ?
It could be better reframed as:
Has any of the compressive sensing pipeline tools (measurement matrices, randomization, L1 or better solvers) allowed one to make a discovery that was not possible before (with the other tools) ?


First, I think this is an unfair question because I don' t feel like it should even be asked! Indeed, nobody asks the obvious question as to whether full rank linear systems and least squares solvers have led to scientific discoveries (they have). On the other hand, a more elaborate technique has to provide an enhanced threshold of justification only a few years after it has been theoretically justified. But here is where it gets weird. I initially indicated that I felt that the following papers from Bruno Ohlshausen were compressive sensing related discoveries:

Emergence of Simple-Cell Receptive Field Properties by Learning a Sparse Code for Natural Images

Olshausen BA, Field DJ (1996).   Nature, 381: 607-609.  reprint (pdf)  |  abstract

Natural Image Statistics and Efficient Coding

Olshausen BA, Field DJ (1996).   Presented at the Workshop on Information Theory and the Brain , September 4-5, 1995, University of Stirling, Scotland. Published in Network, 7: 333-339.   reprint (pdf)  |  abstract


Indeed, after the publication of these papers, the community at large began to realize that sparse coding was not just an artifact of being into the parcimony business. Rather it was an actual biological process that could be mapped to specific cells and a specific area of the brain.

That example did not seem to fit the bill as the paper predated the 2004 papers of Candes, Tao, Romberg and that of Donoho. As such it would not count as compressive sensing.

This was a little disheartening as many people were doing compressive sensing before 2004 (see The invention of compressive sensing) with potentially a link to Prony back to 1796. Further, the clock did not start ticking back in 2004 or 2006, rather it probably began ticking in 2008/2009. Indeed from 2004 till 2007, several measurement matrices allowed nonlinear recovery of sparse signals. In fact during that time frame, there was no technical way of figuring out a simple way whether a specific measurement matrix would allow generic recovery of sparse signals (RIP is NP-Hard to check). It is only in 2007/2008 that generic phase transitions were discovered and eventually we had to wait until 2011 to get even better measurement ensembles beyond strictly random gaussian ensembles. In short, the clock started ticking five years ago, not ten. Given all this background, 

Has there been any discovery or prediction that has been enabled by compressive sensing within the past five years that could not be predicted before ?

I can think of at least two examples:

Compressive ghost imaging by Ori Katz, Yaron Bromberg, and Yaron Silberberg

Why ? Up until that point, ghost imaging was thought to be related to quantum mechanics. Even though various tests were "proving" it was not a quantum mechanical effect, that paper put the last nail to that coffin: The effect is interesting but it ain't quantum mechanical, period.

Applying compressed sensing to genome-wide association studies by Shashaank Vattikuti, James J Lee, Christopher C Chang, Stephen D H Hsu, and Carson C Chow

Why ? because a least squares solver is incapable of enabling a prediction of the type given in that paper. Here thanks to the phase transition found by Tanner and Donoho, one can predict within the linear model how many people are needed to figure out a genetic connection to a specific trait. This is new. You can argue that the linear model is wrong but this is a prediction for that model. There is no similar prediction capibility for a least squares solver.

In the future, I personally think that the map makers are likely to be on the right track to make scientific discoveries.

If you feel that there is a discovery I did not mention, feel free to add your candidate to the comment section of this entry below or in this attendant LinkedIn discussion thread.


 
 
 
Join the CompressiveSensing subreddit or the Google+ Community and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Printfriendly