Friday, November 13, 2015

False Discoveries Occur Early on the Lasso Path - implementation -

A phase transition for the LASSO


False Discoveries Occur Early on the Lasso Path by Weijie Su, Malgorzata Bogdan, Emmanuel Candes

In regression settings where explanatory variables have very low correlations and where there are relatively few e ects each of large magnitude, it is commonly believed that the Lasso shall be able to nd the important variables with few errors|if any. In contrast, this paper shows that this is not the case even when the design variables are stochastically independent. In a regime of linear sparsity, we demonstrate that true features and null features are always interspersed on the Lasso path, and that this phenomenon occurs no matter how strong the e ect sizes are. We derive a sharp asymptotic trade-o between false and true positive rates or, equivalently, between measures of type I and type II errors along the Lasso path. This trade-o states that if we ever want to achieve a type II error (false negative rate) under a given threshold, then anywhere on the Lasso path the type I error (false positive rate) will need to exceed a given threshold so that we can never have both errors at a low level at the same time. Our analysis uses tools from approximate message passing (AMP) theory as well as novel elements to deal with a possibly adaptive selection of the Lasso regularizing parameter. 

The matlab implementation to draw the Lasso Trade-off Diagram is here.

Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Thursday, November 12, 2015

The Big List of Deep Learning Toolkits - implementation -

Much like in the other highly technical reference pages, Kyle McDonald decided to have a big list of Deep Learning toolkits, here it is in Google Docs.




Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

The Human Kernel

Interesting way of doing supervised learning of kernel. From the introduction:

Our work is intended as a preliminary step towards building probabilistic kernel machines that en- capsulate human-like support and inductive biases. Since state of the art machine learning methods perform conspicuously poorly on a number of extrapolation problems which would be easy for humans [10], such efforts have the potential to help automate machine learning and improve perfor- mance on a wide range of tasks – including settings which are difficult for humans to process (e.g., big data and high dimensional problems). Finally, the presented framework can be considered in a more general context, where one wishes to efficiently reverse engineer interpretable properties of any model (e.g., a deep neural network) from its predictions. 

The demos' address are below:

The Human Kernel by Andrew Gordon Wilson, Christoph Dann, Christopher G. Lucas, Eric P. Xing

Bayesian nonparametric models, such as Gaussian processes, provide a compelling framework for automatic statistical modelling: these models have a high degree of flexibility, and automatically calibrated complexity. However, automating human expertise remains elusive; for example, Gaussian processes with standard kernels struggle on function extrapolation problems that are trivial for human learners. In this paper, we create function extrapolation problems and acquire human responses, and then design a kernel learning framework to reverse engineer the inductive biases of human learners across a set of behavioral experiments. We use the learned kernels to gain psychological insights and to extrapolate in human-like ways that go beyond traditional stationary and polynomial kernels. Finally, we investigate Occam's razor in human and Gaussian process based function learning.


Demos of experiments for The Human Kernel can be found here.

 
 
 
Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Wednesday, November 11, 2015

Deep Kernel Learning

Using KISS-GP (see the preprint below), the authors take a look at composing Kernel 'à la CNN'


Deep Kernel Learning by Andrew Gordon Wilson, Zhiting Hu, Ruslan Salakhutdinov, Eric P. Xing

We introduce scalable deep kernels, which combine the structural properties of deep learning architectures with the non-parametric flexibility of kernel methods. Specifically, we transform the inputs of a spectral mixture base kernel with a deep architecture, using local kernel interpolation, inducing points, and structure exploiting (Kronecker and Toeplitz) algebra for a scalable kernel representation. These closed-form kernels can be used as drop-in replacements for standard kernels, with benefits in expressive power and scalability. We jointly learn the properties of these kernels through the marginal likelihood of a Gaussian process. Inference and learning cost O(n) for n training points, and predictions cost O(1) per test point. On a large and diverse collection of applications, including a dataset with 2 million examples, we show improved performance over scalable Gaussian processes with flexible kernel learning models, and stand-alone deep architectures.

Kernel Interpolation for Scalable Structured Gaussian Processes (KISS-GP) by Andrew Gordon Wilson, Hannes Nickisch
We introduce a new structured kernel interpolation (SKI) framework, which generalises and unifies inducing point methods for scalable Gaussian processes (GPs). SKI methods produce kernel approximations for fast computations through kernel interpolation. The SKI framework clarifies how the quality of an inducing point approach depends on the number of inducing (aka interpolation) points, interpolation strategy, and GP covariance kernel. SKI also provides a mechanism to create new scalable kernel methods, through choosing different kernel interpolation strategies. Using SKI, with local cubic kernel interpolation, we introduce KISS-GP, which is 1) more scalable than inducing point alternatives, 2) naturally enables Kronecker and Toeplitz algebra for substantial additional gains in scalability, without requiring any grid data, and 3) can be used for fast and expressive kernel learning. KISS-GP costs O(n) time and storage for GP inference. We evaluate KISS-GP for kernel matrix approximation, kernel learning, and natural sound modelling.
h/t Russ

Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Tuesday, November 10, 2015

TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems - implementation -

Unless you have been hiding in a cave that has no Wifi and 3G, you may have heard of Google releasing TensorFlow. Here is the whitepaper and the attendant video of Jeff Dean at the recent Baylearn meetup. Of note since the release this question on Reddit ( So, should I scrap theano, torch, caffe, and dive into TensorFlow? ), this blog post, and a tip on running it on Amazon EC3.
TensorFlow: Large-Scale Machine Learning on Heterogeneous Distributed Systems by Martın Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dan Mane, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viegas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng

TensorFlow [1] is an interface for expressing machine learning algorithms, and an implementation for executing such algorithms. A computation expressed using TensorFlow can be executed with little or no change on a wide variety of heterogeneous systems, ranging from mobile devices such as phones and tablets up to large-scale distributed systems of hundreds of machines and thousands of computational devices such as GPU cards. The system is flexible and can be used to express a wide variety of algorithms, including training and inference algorithms for deep neural network models, and it has been used for conducting research and for deploying machine learning systems into production across more than a dozen areas of computer science and other fields, including speech recognition, computer vision, robotics, information retrieval, natural language processing, geographic information extraction, and computational drug discovery. This paper describes the TensorFlow interface and an implementation of that interface that we have built at Google. The TensorFlow API and a reference implementation were released as an open-source package under the Apache 2.0 license in November, 2015 and are available at http://www.tensorflow.org


Slides are here.
Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Job: Postdoc positions ERC Advanced Grant A-DATADRIVE-B, KU Leuven Belgium


Marco Signoretto just sent me the following:


Hi Igor,

can you please post this announcement on nuit blanche? thanks!

regards

Marco
Sure Marco, here it is:


Postdoc positions ERC Advanced Grant A-DATADRIVE-B

The research group KU Leuven ESAT-STADIUS is currently offering 2 Postdoc positions (1-year) within the framework of the ERC (European Research Council) Advanced Grant A-DATADRIVE-B (PI: Johan Suykens) http://www.esat.kuleuven.be/stadius/ADB on Advanced Data-Driven Black-box modelling.

The research positions relate to the following possible topics:

-1- Prior knowledge incorporation
-2- Kernels and tensors
-3- Modelling structured dynamical systems
-4- Sparsity
-5- Optimization algorithms
-6- Core models and mathematical foundations
-7- Next generation software tool

The research group ESAT-STADIUS http://www.esat.kuleuven.be/stadius at the university KU Leuven Belgium provides an excellent research environment being active in the broad area of mathematical engineering, including systems and control theory, neural networks and machine learning, nonlinear systems and complex networks, optimization, signal processing, bioinformatics and biomedicine.

The research will be conducted under the supervision of Prof. Johan Suykens. Interested candidates having a solid mathematical background and PhD degree can on-line apply by following the submission guidelines given at the website http://www.esat.kuleuven.be/stadius/ADB/vacancies.php by including CV and motivation letter.




Credit: NASA/Johns Hopkins University Applied Physics Laboratory/Southwest Research Institute
Ice Volcanoes on Pluto?
Release Date: November 9, 2015
Keywords: LORRI, PlutoThe informally named feature Wright Mons, located south of Sputnik Planum on Pluto, is an unusual feature that's about 100 miles (160 kilometers) wide and 13,000 feet (4 kilometers) high. It displays a summit depression (visible in the center of the image) that's approximately 35 miles (56 kilometers) across, with a distinctive hummocky texture on its sides. The rim of the summit depression also shows concentric fracturing. New Horizons scientists believe that this mountain and another, Piccard Mons, could have been formed by the 'cryovolcanic' eruption of ices from beneath Pluto's surface.

Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Monday, November 09, 2015

GURLS vs LIBSVM: Performance Comparison of Kernel Methods for Hyperspectral Image Classification

A while back, the study was made comparing random features to scattering networks features for hyperspectral imagery. This time, the authors look at the difference between using LIBSVM and GURLS (with an implementation of random features) when performing classification in that field.
 



GURLS vs LIBSVM: Performance Comparison of Kernel Methods for Hyperspectral Image Classification by  Nikhila Haridas , V. Sowmya , K. P. Soman 
Kernel based methods have emerged as one of the most promising techniques for Hyper Spectral Image classification and has attracted extensive research efforts in recent years. This paper introduces a new kernel based framework for Hyper Spectral Image (HSI) classification using Grand Unified Regularized Least Squares (GURLS) library. The proposed work compares the performance of different kernel methods available in GURLS package with the library for Support Vector Machines namely, LIBSVM. The assessment is based on HSI classification accuracy measures and computation time. The experiment is performed on two standard Hyper Spectral datasets namely, Salinas A and Indian Pines subset captured by AVIRIS (Airborne Visible Infrared Imaging Spectrometer) sensor. From the analysis, it is observed that GURLS library is competitive to LIBSVM in terms of its prediction accuracy whereas computation time seems to favor LIBSVM. The major advantage of GURLS toolbox over LIBSVM is its simplicity, ease of use, automatic parameter selection and fast training and tuning of multi-class classifier. Moreover, GURLS package is provided with an implementation of Random Kitchen Sink algorithm, which can easily handle high dimensional Hyper Spectral Images at much lower computational cost than LIBSVM.
 
 
Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Guarantees of Riemannian Optimization for Low Rank Matrix Recovery

At last, the use of phase transitions to figure out if manifold based techniques can do well:

Guarantees of Riemannian Optimization for Low Rank Matrix Recovery by Ke Wei, Jian-Feng Cai, Tony F. Chan, Shingyu Leung

We establish theoretical recovery guarantees of a family of Riemannian optimization algorithms for low rank matrix recovery, which is about recovering an $m\times n$ rank $r$ matrix from $p < mn$ number of linear measurements. The algorithms are first interpreted as the iterative hard thresholding algorithms with subspace projections. Then based on this connection, we prove that if the restricted isometry constant $R_{3r}$ of the sensing operator is less than $C_\kappa /\sqrt{r}$ where $C_\kappa$ depends on the condition number of the matrix, the Riemannian gradient descent method and a restarted variant of the Riemannian conjugate gradient method are guaranteed to converge to the measured rank $r$ matrix provided they are initialized by one step hard thresholding. Empirical evaluation shows that the algorithms are able to recover a low rank matrix from nearly the minimum number of measurements necessary.
 
Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Printfriendly