Showing posts with label ELM. Show all posts
Showing posts with label ELM. Show all posts

Tuesday, March 17, 2015

Enhanced Image Classification With a Fast-Learning Shallow Convolutional Neural Network

It must be the second shallow network featured here recently and all use some sort of random projection techniques for their measurements: Enhanced Image Classification With a Fast-Learning Shallow Convolutional Neural Network by Mark D. McDonnell, Tony Vladusich

We present a neural network architecture and training method designed to enable very rapid training and low implementation complexity. Due to its training speed and very few tunable parameters, the method has strong potential for embedded hardware applications requiring frequent retraining or online training. The approach is characterized by (a) convolutional filters based on biologically inspired visual processing filters, (b) randomly-valued classifier-stage input weights, (c) use of least squares regression to train the classifier output weights in a single batch, and (d) linear classifier-stage output units. We demonstrate the efficacy of the method as an image classifier, obtaining state-of-the-art results on the MNIST (0.37% error) and NORB-small (2.2%) image classification databases, with very fast training times compared to standard deep network approaches. The network's performance on the Google Street View House Number (SVHN) (4%) database is also competitive with state-of-the art methods.  
 
 
Join the CompressiveSensing subreddit or the Google+ Community and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Friday, January 09, 2015

Fast, simple and accurate handwritten digit classification using extreme learning machines with shaped input-weights

I first heard about Extreme Learning Machines from a presentation at the Paris Machine Learning Meetup #6 Season 1 on Botnet detection by Joseph Ghafari (a resource for this approach can be found here). These constructions are essentially nonlinear functions of random linear projections. Today, we have a paper that aims at showing how one/few iterations of these functions can provide similar results as deep neural networks.



Fast, simple and accurate handwritten digit classification using extreme learning machines with shaped input-weights by Mark D. McDonnell, Migel D. Tissera, André van Schaik, Jonathan Tapson
Deep networks have inspired a renaissance in neural network use, and are becoming the default option for difficult tasks on large datasets. In this report we show that published deep network results on the MNIST handwritten digit dataset can straightforwardly be replicated (error rates below 1%, without use of any distortions) with shallow 'Extreme Learning Machine' (ELM) networks, with a very rapid training time (~10 minutes). When we used distortions of the training set we obtained error rates below 0.6%. To achieve this performance, we introduce several methods for enhancing ELM implementation, which individually and in combination can significantly improve performance, to the point where it is nearly indistinguishable from deep network performance. The main innovation is to ensure each hidden-unit operates only on a randomly sized and positioned patch of each image. This form of random 'receptive field' sampling of the input ensures the input weight matrix is sparse, with about 90 percent of weights equal to zero, which is a potential advantage for hardware implementations. Furthermore, combining our methods with a small number of iterations of a single-batch backpropagation method can significantly reduce the number of hidden-units required to achieve a particular performance. Our close to state-of-the-art results for MNIST suggest that the ease of use and accuracy of ELM should cause it to be given greater consideration as an alternative to deep networks applied to more challenging datasets.  
 
 
Join the CompressiveSensing subreddit or the Google+ Community and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Saturday, March 15, 2014

Slides: Paris Machine Learning Meetup presentations (Meetup #2 through #9)

We've had 9 meetups so far and have three more to go till the Summer break. Our group is diverse and now counts more than 660 members. While everything is in the archives, I thought a simple list would be clearer, here it is:

Kaggle
Tools/Codes/Implementations

Crowdfunding




Image Credit: NASA/JPL/Space Science Institute
Full-Res: W00087163.jpg was taken on March 13, 2014 and received on Earth March 13, 2014. The camera was pointing toward SATURN at approximately 1,049,429 miles (1,688,892 kilometers) away, and the image was taken using the CB2 and IRP0 filters.

Join the CompressiveSensing subreddit or the Google+ Community and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Monday, December 30, 2013

Video (in French): Machine Learning Meetup #6: Jouer avec Kaggle / Detection de Botnets

The video for the Paris Machine Learning Meetup #6 has been out for a little while now but I just noticed I did not listed here before (all the archives are here) The video (in French) is at:



This 6th edition was about Kaggle and Botnet detection with Neural Networks (Jouer avec Kaggle / Detection de Botnets. The meetup took place on December 11th, 2013 at DojoEvents, a great venue to host meetups. Here are the slides:


A summary is at Paris Machine Learning Meetup #6 Summary and a follow up thought was written in a full blown  Sunday Morning Insight:  Randomization is not a dirty word

One of the speakers was covered by Mathilde Damgé at Le Monde magazine in Kaggle, le site qui transforme le « big data » en or.

 Next Meetup is January 15th, see you there !


Sunday, December 15, 2013

Sunday Morning Insight: Randomization is not a dirty word

From [8]

The recent announcement of Yann LeCun's appointment as a director of the new Artificial Intelligence Lab at Facebook and Geoff Hinton's move with Google sends a pretty clear signal that neural networks are OK and that the time has come for 'AI for images' i.e. image understanding. Beyond the cleverness of the algorithms developed by those two researchers and their collaborators, lies several technical prowesses. One of the most important one is hardware based:  Moore's Law, one of the steamrollers (Predicting the Future: The Steamrollers) that now allows the computation of parameters from larger neural networks with desktop computers.

Parallel to these developments (highlighted in Predicting the Future: Randomness and Parsimony), is the mathematics behind the study of high dimensional objects, a field that has really taken off since compressive sensing back in 2004. In fact, mathematically speaking, the ball got really rolling with Johnson and Lindenstrauss back in 1984. With these results, randomization suddenly became something of a wonder tool. In linear algebra and attendant fields, its use is only starting to get noticed (see the recent Slowly but surely they'll join our side of the Force... ). In fact, several folks decided to give a word to this field: Randomized Numerical Linear Algebra (RandNLA) -Most blog entries on the subject are under the RandNLA tag)- 

All is well and peachy keen but how is randomization connected to the nonlinear world of deep learning and neural networks ?

If you were to read the literature, you'd think you'd be watching some episode of Game of Thrones. Not the treacherous and gratitious violence, but rather the many tribes living in their own islands of knowledge. Here is a small case in point: as Joseph Ghafari was making a presentation at the recent Paris Machine Learning Meetup #6 on designing neural networks for botnet detection, I was struck by the similarity between the ELM approach to neural networks spearheaded by Guang-Bin Huang since 2000 [22] and the work on Random Kitchen Sinks [1,2,13] or the work on 1bit and quantization CS [23,24,25] and even more recent work on high dimensional function estimation with ridge functionals [3-9]

Similarly, I do not see papers proposing SVM with Extreme Learning Machine (ELM) techniques mentioning the random kitchen sinks approach,but I am sure I am not reading the right papers either.




None of this really matter. It is a question of time before they actually do talk to each other. But much like what happened in compressive sensing where we had more than 50 sparsity seeking solvers, we really had no real way of judging which one was really interesting. The sharp phase transitions ( see Sunday Morning Insight: The Map Makers ) eventually were the only way to figure this one out. One wonders if they could provide rules of thumbs for understanding these nonlinear approximations of the identity [14]. At the very least, much like in Game of Thrones, the walls, or the sharp phase transition will act as an acid test for all involved. Here are some of the questions that could be answered. From the ELM theory Open Problems page:


  1. As observed in experimental studies, the performance of basic ELM is stable in a wide range of number of hidden nodes. Compared to the BP learning algorithm, the performance of basic ELM is not very sensitive to the number of hidden nodes. However, how to prove it in theory remains open.
  2. One of the typical implementations of ELM is to use random nodes in the hidden layer and the hidden layer of SLFNs need not be tuned. It is interesting to see that the generalization performance of ELM turns out to be very stable. How to estimate the oscillation bound of the generalization performance of ELM remains open too.
  3. It seems that ELM performs better than other conventional learning algorithms in applications with higher noise. How to prove it in theory is not clear.
  4. ELM always has faster learning speed than LS-SVM if the same kernel is used?
  5. ELM provides a batch learning kernel solution which is much simpler than other kernel learning algorithms such as LS-SVM. It is known that it may not be straightforward to have an efficient online sequential implementation of SVM and LS-SVM. However, due to the simplicity of ELM, is it possible to implement the online sequential variant of the kernel based ELM?
  6. ELM always provides similar or better generalization performance than SVM and LS-SVM if the same kernel is used (if not affected by computing devices' precision)?
  7. ELM tends to achieve better performance than SVM and LS-SVM in multiclasses applications, the higher the number of classes is, the larger the difference of their generalization performance will be?
  8. Scalability of ELM with kernels in super large applications.
  9. Parallel and distributed computing of ELM.
  10. ELM will make real-time reasoning feasible?
  11. The hidden node / neuron parameters are not only independent of the training data but also of each other.
  12. Unlike conventional learning methods which MUST see the training data before generating the hidden node / neuron parameters, ELM could randomly generate the hidden node / neuron parameters before seeing the training data.
So far only few works seem to find those transitions. To a certain extent, it looks to be the case in the quantization compressive sensing case [23]  

From [23]


From [8]

or in the examples featured in last  Sunday Morning Insight's that talked about Sharp Phase Transitions in Machine Learning.. . in all, while randomization helps those nonlinear models in getting better results, the acid test that will allow us figure out which one is the most robust, will come from their behaviors with respect the sharp phase transitions...If you are presently using a specific algorithm: Are you going to provide the rest of the community with new charts ? Are you going to be the next Map Makers ?


References:
[2] Random Features for Large-Scale Kernel Machines by Ali Rahimi and Ben Recht
We consider the problem of learning multi-ridge functions of the form f(x) = g(Ax) from point evaluations of f. We assume that the function f is defined on an l_2-ball in R^d, g is twice continuously differentiable almost everywhere, and A \in R^{k \times d} is a rank k matrix, where k << d. We propose a randomized, polynomial-complexity sampling scheme for estimating such functions. Our theoretical developments leverage recent techniques from low rank matrix recovery, which enables us to derive a polynomial time estimator of the function f along with uniform approximation guarantees. We prove that our scheme can also be applied for learning functions of the form: f(x) = \sum_{i=1}^{k} g_i(a_i^T x), provided f satisfies certain smoothness conditions in a neighborhood around the origin. We also characterize the noise robustness of the scheme. Finally, we present numerical examples to illustrate the theoretical bounds in action.


Let us assume that f is a continuous function defined on the unit ball of Rd, of the form f(x)=g(Ax), where A is a k×d matrix and g is a function of k variables for k≪d. We are given a budget m∈N of possible point evaluations f(xi), i=1,...,m, of f, which we are allowed to query in order to construct a uniform approximating function. Under certain smoothness and variation assumptions on the function g, and an {\it arbitrary} choice of the matrix A, we present in this paper
1. a sampling choice of the points {xi} drawn at random for each function approximation;
2. algorithms (Algorithm 1 and Algorithm 2) for computing the approximating function, whose complexity is at most polynomial in the dimension d and in the number m of points.
Due to the arbitrariness of A, the choice of the sampling points will be according to suitable random distributions and our results hold with overwhelming probability. Our approach uses tools taken from the {\it compressed sensing} framework, recent Chernoff bounds for sums of positive-semidefinite matrices, and classical stability bounds for invariant subspaces of singular value decompositions.

We study properties of ridge functions $f(x)=g(a\cdot x)$ in high dimensions $d$ from the viewpoint of approximation theory. The considered function classes consist of ridge functions such that the profile $g$ is a member of a univariate Lipschitz class with smoothness $\alpha > 0$ (including infinite smoothness), and the ridge direction $a$ has $p$-norm $\|a\|_p \leq 1$. First, we investigate entropy numbers in order to quantify the compactness of these ridge function classes in $L_{\infty}$. We show that they are essentially as compact as the class of univariate Lipschitz functions. Second, we examine sampling numbers and face two extreme cases. In case $p=2$, sampling ridge functions on the Euclidean unit ball faces the curse of dimensionality. It is thus as difficult as sampling general multivariate Lipschitz functions, a result in sharp contrast to the result on entropy numbers. When we additionally assume that all feasible profiles have a first derivative uniformly bounded away from zero in the origin, then the complexity of sampling ridge functions reduces drastically to the complexity of sampling univariate Lipschitz functions. In between, the sampling problem's degree of difficulty varies, depending on the values of $\alpha$ and $p$. Surprisingly, we see almost the entire hierarchy of tractability levels as introduced in the recent monographs by Novak and Wo\'zniakowski.
[10] Conic Geometric Programming by Venkat Chandrasekaran and Parikshit Shah
[12] Random Conic Pursuit for Semidefinite Programming by Ariel Kleiner, Ali Rahimi, Michael I. Jordan
We present a novel algorithm, Random Conic Pursuit, that solves semidefinite programs (SDPs) repeatedly solving randomly generated optimization problems over two-dimensional subcones of the PSD cone. This scheme is simple, easy to implement, applicable to general nonlinear SDPs, and scalable. These advantages come at the expense of inexact solutions, though we show these to be practically useful and provide theoretical guarantees for some of them. This property renders Random Conic Pursuit of particular interest for machine learning applications, in which the relevant SDPs are based on random data and so exact minima are often not a priority.
[13] Random Features for Large-Scale Kernel Machines by Ali Rahimi and Ben Recht
[17] Benoît Frénay's blog , http://bfrenay.wordpress.com/extreme-learning/
[21] Nystrom Method vs Random Fourier Features:: A Theoretical and Empirical Comparison  by Tianbao Yang, Yu-Feng Li, Mehrdad Mahdavi, Rong Jin, Zhi-Hua Zhou
[22] Extreme Learning Machine (ELM), http://www.ntu.edu.sg/home/egbhuang/
[23] J. N. Laska Regime Change. (includes: BSE, several algorithms, numerical work, and regime change)
[24] http://dsp.rice.edu/1bitCS/

Thursday, December 12, 2013

Paris Machine Learning Meetup #6 Summary

[updated 12/12/13, 10PM]

So last night, we had our Paris Machine Learning meetup (#6). Our group has more than 425 members which probably makes it one the largest machine learning meetup in Europe (second behind London). Our LinkedIn group has 199 members and already several job offers there (please post them there with the indication [job] in the title ). Next meeting is January 15th.

The people counter on Meetup indicated 111 people. I counted 65 physically in the room. The counter went up to a whopping 137 people two days before the meetup. We were in some  'attention competition' with two other interesting meetups (node.js and Leweb13). At least, we did not get train strikes like today. Again, Franck and I are enormously grateful to the folks who changed their RSVPs.

DojoEvents hosted us once again. Their value proposition is listed here. If you are a startup and want physical hosting in Paris, check them out. They also have an event based outfit if you want to organize meetups, hackatons and so forth. Guillaume Pellerin used his Telecasting solution from Parisson to do the streaming and the recording of the meeting. 
Matthieu provided a good overview of what he and his team at Dataiku are doing in the Yandex Kaggle challenge. He also provided the code for parsing that competition's dataset. It is available at:


The Dataiku team (currently ranked 11th on the Yandex Challenge) is also interested in people joining them to improve their scores. During the presentation, Matthieu mentioned the long training time but I felt it probably was an instance of low rank matrix completion if one sets it up the right way [4]. All in all, a very nice introduction in the subject and a strong technical Q&A like we like them.
On a different note, Dataiku is also hiring.


[ Mathilde Damgé of Le Monde.fr wrote about Matthieu et tangentially on the Meetup in Kaggle, le site qui transforme le « big data » en or]

Joseph showed us the design of the ELM approach to neural networks he used on a botnet at OrangeLabs with Stephane. More on that later. Suffice to say, those ELM networks are very similar to the Random Kitchen Sinks/FastFood implementation as well as other works (see also the Summer of Deeper Kernels) that I will probably feature in this upcoming Sunday Morning Insight. I initially was expecting a run-of-the-mill neural network presentation, instead I got blown away: 14 seconds of training, wow. You know what ? maybe RKS/ELM should be tried on the Yandex set...I am just saying.

In order to provide some context, Franck and I inserted a What's New presentation, I think it was well received. In particular, we mentioned the NIPS announcements as well as the PSIM program as Machine Learning is always a crosscutting technology and that it really has the potential to provide a very large ROI:
LeCun's announcement at Facebook and Geoff Hinton's move with Google sends a pretty clear signal  that neural networks are OK and that the time has come for 'AI for images' i.e. image understanding. In my mind, since the both of them are in the neural network vein, it is also time we get to figure out how those neural network hit the phase transition and how this provides us with rules of thumbs for designing them see[1,2]. Gabriel, one the meetuper who was at NIPS, wrote a small summary of what happened at the announcement.

Thank you to both of our speakers Matthieu Scordia and Joseph Ghafari who were good sports in answering our, sometimes, long winded questions.

In the champagne drinking part of the meeting, we discussed education analytics (we need to talk about SPARFA and similar solvers), how to get Danny Bickson to give a talk in Paris on GraphLab (if you are somewhere else in Europe and you -or better your company- have an interest in him coming across the pond, please email and let's talk), but also how the french program PSIM could potentially be a fit to businesses that are just on the right side of impossible [3] (That would include applications of this). 

Live Streaming was provided during the presentations and a recording should be online shortly.

Wednesday, December 11, 2013

Paris Machine Learning Meetup #6: Playing with Kaggle/ Botnet detection with Neural Networks

Tonight starting a 7:00 pm local Paris time, we'll have our 6th Paris Machine Learning meetup featuring Matthieu Scordia  (Playing with Kaggle) and Joseph Ghafari (Botnet detection with Neural Networks). It'll be held at DojoEvents> The twitter handle is #MLParis
Live streaming will start between 7:30pm and 7:45pm and will be provided by the Telecasting solution from Parisson. The presentations are likely to be in French with slides in English.

We have about 120 people registered but we expect that number to drop. 

We also have the video of the previous meetup (in French) held on November 13, 2013 also at DojoEvents that was focused on Making Sense of Two Data Tsunamis: Genomics and the Internet of Things / La Génomique et l'Internet des Objets. The slide presentations are listed below:

Friday, November 22, 2013

Paris Machine Learning Meetup #6: Playing with Kaggle/ Botnet detection with Neural Networks. Jouer avec Kaggle / Detection de Botnets

Before I get to the announcement for the next Paris Machine Learning meetup #6, let me point to a meetup that will take place in about 9 hours in Silicon Valley on Compressed sensing techniques for sensor data using unsupervised learning by Song Cui. Coming back to the focus of this announcement, on December 11th, we will have the sixth Paris Machine Learning Meetup: Playing with Kaggle/ Botnet detection with Neural Networks. Jouer avec Kaggle / Detection de Botnets. Here is the announcement in French.

Mercredi 11 Decembre à 19h00 à DojoEvents, 41 Boulevard Saint-Martin, Paris 

Pour ce sixieme rendez-vous du Machine Learning à Paris nous aurons au moins deux présentateurs: Matthieu Scordia (Profil Kaggle: http://www.kaggle.com/users/70112/matt-sco ) et Joseph Ghafari ( http://www.josephghafari.com/ ) . Matthieu est classe 147eme sur 130,000 utilisateurs de Kaggle. Il nous parlera des différentes stratégies mises en place pour concourir dans les differents challenges de la plateforme Kaggle. Joseph, lui, nous parlera de l'utilisation d'un reseau de neurones pour la detection de bots sur les reseaux d'un exploitant telecom francais. Les présentations ainsi que le Champagne seront en français.

Résumé des présentations:

* Les compétitions Kaggle, un moyen fun et instructif pour mesurer ses compétences en machine learning. Matthieu nous invitera à entrer dans un challenge et donnera quelques astuces/pièges à éviter. Matthieu Scordia (profil Kaggle: http://www.kaggle.com/users/70112/matt-sco ), sa bio: Master's degree in Artificial Intelligence @ UPMC.
Data Scientist Junior @ Dataiku.

* "Réseaux de neurones pour la détection de Botnets". Le but de cette etude est d'appliquer de nouveaux modèles de réseaux de neurones pour classifier des traces de trafic internet issues de serveurs DNS d'Orange. Un apprentissage supervisé est appliqué à ces données pour : détecter les traces issues de trafic malveillant (Botnets en particulier) et déterminer les paramètres réseaux les plus pertinents pour caractériser le passage d'un Bot sur le réseau. Joseph Ghafari ( http://www.josephghafari.com/ )

Si vous ou un collègue, voulez-vous inscrire a la liste du meetup, vous pouvez le faire directement à

http://www.meetup.com/Paris-Machine-learning-applications-group/

Pour s'inscrire au meetup #6 c'est ici: http://www.meetup.com/Paris-Machine-learning-applications-group/events/150851882/

Les archives des précédents meetups se trouvent ici:
http://nuit-blanche.blogspot.com/p/paris-based-meetups-on-machine-learning.html

Il y aussi un groupe Paris Machine Learning sur LinkedIn:
http://www.linkedin.com/groups?gid=6400776

Les organisateurs,

Franck Bardol, Frederick Demback, Igor Carron
 


Join the CompressiveSensing subreddit or the Google+ Community and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Printfriendly