Showing posts with label ICLR2015. Show all posts
Showing posts with label ICLR2015. Show all posts

Friday, January 23, 2015

In Search of the Real Inductive Bias: On the Role of Implicit Regularization in Deep Learning

How could I have missed this one from the papers currently in review for ICLR 2015 ? In fact, I did not miss it, I read it and then ... other things took over. So without further ado, here is the starting point of the study:

...Consider, however, the results shown in Figure 1, where we trained networks of increasing size on the MNIST and CIFAR-10 datasets. Training was done using stochastic gradient descent with momentum and diminishing step sizes, on the training error and without any explicit regularization. As expected, both training and test error initially decrease. More surprising is that if we increase the size of the network past the size required to achieve zero training error, the test error continues decreasing! This behavior is not at all predicted by, and even contrary to, viewing learning as fitting a hypothesis class controlled by network size...

and then they use advanced matrix factorization to understand the issue better, what's not to like?
 


In Search of the Real Inductive Bias: On the Role of Implicit Regularization in Deep Learning by Behnam Neyshabur, Ryota Tomioka, Nathan Srebro.

We present experiments demonstrating that some other form of capacity control, different from network size, plays a central role in learning multilayer feed-forward networks. We argue, partially through analogy to matrix factorization, that this is an inductive bias that can help shed light on deep learning.
 
Join the CompressiveSensing subreddit or the Google+ Community and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Monday, January 19, 2015

Speeding-up Convolutional Neural Networks Using Fine-tuned CP-Decomposition

In "The papers for ICLR 2015 are now open for discussion !" I mentioned a few papers that Nuit Blanche had featured recently and that were up for review at ICLR 2015. Here is another one that aims at using tensor reduction for reducing the time it takes to perform computation with convolutional neural networks:

Speeding-up Convolutional Neural Networks Using Fine-tuned CP-Decomposition by Vadim Lebedev, Yaroslav Ganin, Maksim Rakhuba, Ivan Oseledets, Victor Lempitsky

We propose a simple two-step approach for speeding up convolution layers within large convolutional neural networks based on tensor decomposition and discriminative fine-tuning. Given a layer, we use non-linear least squares to compute a low-rank CP-decomposition of the 4D convolution kernel tensor into a sum of a small number of rank-one tensors. At the second step, this decomposition is used to replace the original convolutional layer with a sequence of four convolutional layers with small kernels. After such replacement, the entire network is fine-tuned on the training data using standard backpropagation process.
We evaluate this approach on two CNNs and show that it yields larger CPU speedups at the cost of lower accuracy drops compared to previous approaches. For the 36-class character classification CNN, our approach obtains a 8.5x CPU speedup of the whole network with only minor accuracy drop (1% from 91% to 90%). For the standard ImageNet architecture (AlexNet), the approach speeds up the second convolution layer by a factor of 4x at the cost of 1% increase of the overall top-5 classification error.
 
 
Join the CompressiveSensing subreddit or the Google+ Community and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Friday, January 16, 2015

The papers for ICLR 2015 are now open for discussion !

 
 
Seven papers that were recently featured here on Nuit Blanche have also been submitted to ICLR 2015 (they are all under the ICLR2015 tag). 


Others are listed here. 

If you are interested in commenting on them, here is how to go about it. From Hugo Larochelle Google Plus feed:

The papers for ICLR 2015 are now open for discussion!
(this requires an ICLR 2015 CMT account, which you can sign up for if you don't have one already...)
and from here:


The open commenting period for ICLR 2015 has begun.

This year we will be using CMT's public commenting mechanism, so to participate you must log in to the ICLR 2015 CMT site:

https://cmt.research.microsoft.com/ICLR2015/Protected/PublicComment.aspx

If you don't already have a CMT account, please request one.

The list of submitted papers, with links to arXiv, is available at

http://www.iclr.cc/doku.php?id=iclr2015:main

Please remember that comments will be visible to anyone who logs in to CMT. We intend to make the comments available on openreview.net once they finish a major revision of their infrastructure. The anonymous reviews will be visible on CMT at the beginning of the discussion period, and will also eventually be made available on openreview.net.

If you want to know what is meant by open commenting, please see this description of the ICLR reviewing model: http://www.iclr.cc/doku.php?id=pubmodel

If you have any questions, please send them to iclr2015.programchairs@gmail.com 

 
 
 
 
Join the CompressiveSensing subreddit or the Google+ Community and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Monday, January 12, 2015

Improving approximate RPCA with a k-sparsity prior

Using a regularizer borrowed from k-sparse autoencoders, this is interesting:
 
Improving approximate RPCA with a k-sparsity prior by Maximilian Karl, Christian Osendorfer

A process centric view of robust PCA (RPCA) allows its fast approximate implementation based on a special form o a deep neural network with weights shared across all layers. However, empirically this fast approximation to RPCA fails to find representations that are parsemonious. We resolve these bad local minima by relaxing the elementwise L1 and L2 priors and instead utilize a structure inducing k-sparsity prior. In a discriminative classification task the newly learned representations outperform these from the original approximate RPCA formulation significantly.
 
 
Join the CompressiveSensing subreddit or the Google+ Community and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Wednesday, January 07, 2015

Why does Deep Learning work? - A perspective from Group Theory


h/T to Giuseppe Guissepe, Gabriel and Suresh for the discussion.

Here is a new insight from a paper that is up for review at ICLR2015:
We give an informal description. Suppose G is a group that acts on a set X by moving its points around ( e.g., groups of 2 x 2 invertible matrices acting over the euclidean plane). Consider x2X, and let Ox be the set of all points reachable from x via the group action . Ox is called an orbit.2. Some of the group elements may leave x unchanged. This subset,Sx, which is also a subgroup, is the stabilizer of x....
 
To answer, imagine that the space of the autoencoders form a group. A batch of learning iterations, aka search for stabilizers, stops whenever a stabilizer is found. Roughly speaking, if the search is a Markov chain (or a guided chain such as MCMC), then the bigger a stabilizer, the earlier it will be hit. The group structure implies that this big stabilizer corresponds to a small orbit. Now, intuition suggests that the simpler a feature, the smaller is its orbit. For example, a line-segment generates many fewer possible shapes 3 under linear deformations, than a flower-like shape. An autoencoder then should learn these simpler features first, which falls in line with most experiments (see Leeet al. (2009))
This points directly to a new kind of regularizer ! The next Sunday Morning Insight will probably be on the new regularizers. Without futher ado, here is: Why does Deep Learning work? - A perspective from Group Theory by Arnab Paul, Suresh Venkatasubramanian

Why does Deep Learning work? What representations does it capture? How do higher-order representations emerge? We study these questions from the perspective of group theory, thereby opening a new approach towards a theory of Deep learning.
One factor behind the recent resurgence of the subject is a key algorithmic step called pre-training: first search for a good generative model for the input samples, and repeat the process one layer at a time. We show deeper implications of this simple principle, by establishing a connection with the interplay of orbits and stabilizers of group actions. Although the neural networks themselves may not form groups, we show the existence of {\em shadow} groups whose elements serve as close approximations.
Over the shadow groups, the pre-training step, originally introduced as a mechanism to better initialize a network, becomes equivalent to a search for features with minimal orbits. Intuitively, these features are in a way the {\em simplest}. Which explains why a deep learning network learns simple features first. Next, we show how the same principle, when repeated in the deeper layers, can capture higher order representations, and why representation complexity increases as the layers get deeper.

 Let us note that it mentions another paper featured recently in Sunday Morning Insight: An exact mapping between the Variational Renormalization Group and Deep Learning. the authors say the following about that paper:
Mehta & Schwab (2014) recently showed an intriguing connection between Renormalization group flow 5 and deep-learning. They constructed an explicit mapping from a renormalization group over a block-spin Ising model (as proposed by Kadanoff et al. (1976)), to a DL architecture. On the face of it, this result is complementary to ours, albeit in a slightly different settings. Renormalization is a process of coarse-graining a system by first throwing away small details from its model, and then examining the new system under the simplified model (see Cardy (1996)). In that sense the orbit-stabilizer principle is a re-normalizable theory - it allows for the exact same coarse-graining operation at every layer - namely, keeping only minimal orbit shapes and then passing them as new parameters for the next layer - and the theory remains unchanged at every scale.
As we mentioned then, other approaches such as ScatNet also tries to learn a nonlinear relationship (in their words a scattering transform) so that eventual nonlinear features can be used in a group theoretic framework.

 
Join the CompressiveSensing subreddit or the Google+ Community and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Monday, January 05, 2015

Compressing Deep Convolutional Networks using Vector Quantization


Deep convolutional neural networks (CNN) has become the most promising method for object recognition, repeatedly demonstrating record breaking results for image classification and object detection in recent years. However, a very deep CNN generally involves many layers with millions of parameters, making the storage of the network model to be extremely large. This prohibits the usage of deep CNNs on resource limited hardware, especially cell phones or other embedded devices. In this paper, we tackle this model storage issue by investigating information theoretical vector quantization methods for compressing the parameters of CNNs. In particular, we have found in terms of compressing the most storage demanding dense connected layers, vector quantization methods have a clear gain over existing matrix factorization methods. Simply applying k-means clustering to the weights or conducting product quantization can lead to a very good balance between model size and recognition accuracy. For the 1000-category classification task in the ImageNet challenge, we are able to achieve 16-24 times compression of the network with only 1% loss of classification accuracy using the state-of-the-art CNN.
 
 
Join the CompressiveSensing subreddit or the Google+ Community and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Friday, December 26, 2014

Unsupervised Learning of Spatiotemporally Coherent Metrics

 
 From the paper:
 
Non-linear operators consisting of a redundant linear transformation followed by a point-wise non-linearity and a local pooling, are fundamental building blocks in deep convolutional networks. This is due to their capacity to generate local invariance while preserving discriminative information Le-Cun et al. (1998); Bruna & Mallat (2013).
 
 
Unsupervised Learning of Spatiotemporally Coherent Metrics by Ross Goroshin, Joan Bruna, Jonathan Tompson, David Eigen, Yann LeCun

Current state-of-the-art object detection and recognition algorithms rely on supervised training, and most benchmark datasets contain only static images. In this work we study feature learning in the context of temporally coherent video data. We focus on training convolutional features on unlabeled video data, using only the assumption that adjacent video frames contain semantically similar information. This assumption is exploited to train a convolutional pooling auto-encoder regularized by slowness and sparsity. We show that nearest neighbors of a query frame in the learned feature space are more likely to be its temporal neighbors.  
 
 
Join the CompressiveSensing subreddit or the Google+ Community and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Thursday, December 25, 2014

A la Carte - Learning Fast Kernels

Here is a specific tweak to FastFood in order to enable a large set of kernels to better fit group theoretic transformation. It is a little bit a similar path taken in the ScatNet approach.


A la Carte - Learning Fast Kernels by Zichao Yang, Alexander J. Smola, Le Song, Andrew Gordon Wilson

Kernel methods have great promise for learning rich statistical representations of large modern datasets. However, compared to neural networks, kernel methods have been perceived as lacking in scalability and flexibility. We introduce a family of fast, flexible, lightly parametrized and general purpose kernel learning methods, derived from Fastfood basis function expansions. We provide mechanisms to learn the properties of groups of spectral frequencies in these expansions, which require only O(mlogd) time and O(m) memory, for m basis functions and d input dimensions. We show that the proposed methods can learn a wide class of kernels, outperforming the alternatives in accuracy, speed, and memory consumption.  
 
 
Join the CompressiveSensing subreddit or the Google+ Community and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Wednesday, December 24, 2014

On the Stability of Deep Networks

Wow! after this, this and this, we are now converging:

On the Stability of Deep Networks by Raja Giryes, Guillermo Sapiro, Alex M. Bronstein

In this work we study the properties of deep neural networks with random weights. We formally prove that these networks perform a distance-preserving embedding of the data. Based on this we then draw conclusions on the size of the training data and the networks' structure.
 
 
Join the CompressiveSensing subreddit or the Google+ Community and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Tuesday, December 23, 2014

Deep Fried Convnets

Very interesting. Thanks for this summary. I loved the paper of Jimmy Ba and Rich Caruana. Here's one of our recent takes on the issue of kernels, randomization and new forms of deep layerd: http://arxiv.org/abs/1412.7149

In all these, I think it's important to analyse performance on large datasets such as imagenet.
Thanks Nando !  another piece of the puzzle for the construction of reconstruction solvers:



The fully connected layers of a deep convolutional neural network typically contain over 90% of the network parameters, and consume the majority of the memory required to store the network parameters. Reducing the number of parameters while preserving essentially the same predictive performance is critically important for operating deep neural networks in memory constrained environments such as GPUs or embedded devices.
In this paper we show how kernel methods, in particular a single Fastfood layer, can be used to replace all fully connected layers in a deep convolutional neural network. This novel Fastfood layer is also end-to-end trainable in conjunction with convolutional layers, allowing us to combine them into a new architecture, named deep fried convolutional networks, which substantially reduces the memory footprint of convolutional networks trained on MNIST and ImageNet with no drop in predictive performance.

 
 
Join the CompressiveSensing subreddit or the Google+ Community and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Saturday, December 20, 2014

Sunday Morning Insight: The Regularization Architecture


Or Not ?


Today's insights go in different directions. The first one came in large part from a remark by John Platt on  Technet's Machine Learning blog who puts this into words much better than I ever could.
...Given the successes of deep learning, researchers are trying to understand how they work. Ba and Caruana had a NIPS paper which showed that, once a deep network is trained, a shallow network can learn the same function from the outputs of the deep network. The shallow network can’t learn the same function directly from the data. This indicates that deep learning could be an optimization/learning trick...
 
My emphasis on the last sentence. So quite clearly, the model reduction of Jimmy Ba and Rich Caruana that was featured a year ago (Do Deep Nets Really Need to be Deep?, NIPS version Do Deep Nets Really Need to be Deep?) or the more recently featured shallow model  (Kernel Methods Match Deep Neural Networks on TIMIT) do in fact point to a potentially better regularization scheme that we have not really found.
If in fact, one wants to continue to think in terms of several layers, then one could think of these stacked networks as iterations of a reconstruction solver as Christian Schuler, Michael Hirsch, Stefan Harmeling, and Bernhard Scholkopf do in Learning to Deblur.
 
Finally, in yesterday's videos (Saturday Morning Videos : Semidefinite Optimization, Approximation and Applications (Simons Institute @ Berkeley)), one could watch Sanjeev Arora [3] talk about Adventures in Linear Algebra++ and Unsupervised Learning (slides). It so happens that what he describes as Linear Algebra ++ is none other than the Advanced Matrix Factorization Jungle. But more importantly, he mentions that randomly wired deep nets (Provable Bounds for Learning Some Deep Representations)


were acknowledged in the building of a recent deeper network paper [1]. From [1]:

In general, one can view the Inception model as a logical culmination of [12] while taking inspiration and guidance from the theoretical work by Arora et al [2]. The benefits of the architecture are experimentally verified on the ILSVRC 2014 classification and detection challenges, on which it significantly outperforms the current state of the art.
and in the conclusion of that same paper:

Although it is expected that similar quality of result can be achieved by much more expensive networks of similar depth and width, our approach yields solid evidence that moving to sparser architectures is feasible and useful idea in general. This suggest promising future work towards creating sparser and more refined structures in automated ways on the basis of [2].
Deeper constructs yes, but sparser ones.
References:


[1] Going Deeper with Convolutions by Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, Andrew Rabinovich

We propose a deep convolutional neural network architecture codenamed "Inception", which was responsible for setting the new state of the art for classification and detection in the ImageNet Large-Scale Visual Recognition Challenge 2014 (ILSVRC 2014). The main hallmark of this architecture is the improved utilization of the computing resources inside the network. This was achieved by a carefully crafted design that allows for increasing the depth and width of the network while keeping the computational budget constant. To optimize quality, the architectural decisions were based on the Hebbian principle and the intuition of multi-scale processing. One particular incarnation used in our submission for ILSVRC 2014 is called GoogLeNet, a 22 layers deep network, the quality of which is assessed in the context of classification and detection.

[2] Learning to Deblur by Christian J. Schuler, Michael Hirsch, Stefan Harmeling, and Bernhard Scholkopf

We show that a neural network can be trained to blindly deblur images. To accomplish that, we apply a deep layered architecture, parts of which are borrowed from recent work on neural network learning, and parts of which incorporate computations that are specific to image deconvolution. The system is trained end-to-end on a set of artificially generated training examples, enabling competitive performance in blind deconvolution, both with respect to quality and runtime.


 
 
 
 
Join the CompressiveSensing subreddit or the Google+ Community and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Printfriendly