Tuesday, December 27, 2016

Slides: Annual Paris-Saclay Center for Data Science Pitching Day

Here is an interesting initiative: Data Science for Science. During a discussion, Balazs reminded me of the very interesting Annual CDS Pitching Day that occured last month. It was a "meeting... to officially launch the Paris-Saclay Center for Data Science 2.0." The CDS ".. gather(s) data providers and data analysts around the common theme of Data Science. There will be five introductory talks presenting the tools and platforms built by the CDS, and 13 talks given by members of the CDS, covering a wide spectrum of topics on both domain sciences and data science. The aim is to draw an overall picture about the demand for CDS tools and expertise in Saclay and to help PIs building their projects.  "
 
9 novembre 2016
  • 08:30 - 09:30 Welcome coffee
  • 09:30 - 10:45 CDS tools and platforms
    Responsable: Dr. Sarah COHEN-BOULAKIA (Laboratoire de Recherche en Informatique)
    • 09:30 Introduction 15'
      Intervenant: Balázs Kégl (LAL)
      Documents: Slides pdf file
    • 09:45 RAMPs and Data Challenges 15'
      Intervenant: Balázs Kégl (LAL)
      Documents: Slides pdf file
    • 10:00 The linked data platform 15'
      Intervenant: Mrs. Karima Rafes (LRI)
      Documents: Slides pdf file more information link
    • 10:15 The open software initiative 15'
      Intervenant: Alexandre Gramfort (Telecom ParisTech, CNRS)
      Documents: Slides pdf file
    • 10:30 Ultrawalls and interactive visualization 15'
      Intervenant: Emmanuel Pietriga (INRIA)
      Documents: Slides pdf file summary pdf file
  • 10:45 - 11:15 Coffee break
  • 11:15 - 13:05 Life sciences
    Responsable: Alexandre Gramfort (Telecom ParisTech, CNRS)
    • 11:15 Apprentissage automatique pour une aide au codage PMSI 22'
      Intervenant: Namik Taright (APHP)
      Documents: Slides pdf file summary pdf file
    • 11:37 Managing and analyzing analytical chemistry data sets 22'
      Intervenant: Ali Tfayli (LIPSYS / UPSud)
      Documents: Slides pdf file summary pdf file
    • 11:59 Secure cytotoxic administration by non-invasive spectroscopy 22'
      Intervenant: Laetitia Le (LIPSYS / UPSud)
      Documents: Slides link summary pdf file
    • 12:21 Efficiently Ranking big biological and biomedical data sets using rank aggregation techniques 22'
      Intervenant: Dr. Sarah COHEN-BOULAKIA (Laboratoire de Recherche en Informatique)
      Documents: Slides pdf file summary pdf file
    • 12:43 Reconstructing the past: deep learning for population genetics 22'
      Intervenant: Guillaume Charpiat (TAO Team/ INRIA)
      Documents: Transparents pdf file summary pdf file
  • 13:05 - 14:05 Lunch
  • 14:05 - 15:35 Tools and social sciences
    Responsable: Dr. Gael Varoquaux (INRIA)
    • 14:05 Causal discovery for CDS practitioners 22'
      Intervenant: Ms. Isabelle Guyon (UPSud / INRIA / ChaLearn)
      Documents: Slides powerpoint file summary pdf file
    • 14:27 Multi-level data fusion 22'
      Intervenant: Dr. Kaouthar Benameur (ONERA)
      Documents: Slides powerpoint file summary pdf file
    • 14:49 Global scientific positioning system for CDS 22'
      Intervenant: Philippe Caillou (LRI / UPSud)
      Documents: Slides pdf file summary pdf file
    • 15:11 Heterogeneous social data management and mining 22'
      Intervenant: Nacéra Bennacer (LRI)
      Documents: Paper pdf file Slides link summary pdf file
  • 15:35 - 16:05 Coffee break
  • 16:05 - 17:35 Physical sciences
    Responsable: Ms. Isabelle Guyon (UPSud / INRIA / ChaLearn)
    • 16:05 A platform for decoding and recoding turbulence 22'
      Intervenant: Dr. Bérengère Podvin (CNRS)
      Documents: Slides pdf file summary pdf file
    • 16:27 A framework for sensor data management and analysis 22'
      Intervenant: Karine Zeitouni (DAVID / UVSQ)
      Documents: Slides pdf file summary pdf file
    • 16:49 Generic stereoscopic tools for planetary topography 22'
      Intervenant: Mr. Frédéric Schmidt (GEOPS / UPSud)
      Documents: Slides pdf file summary pdf file
    • 17:11 Tracking machine learning challenge 22'
      Intervenant: Dr. David Rousseau (ATLAS)
      Documents: Transparents pdf file summary pdf file
 
 
 
 
 
 
Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Job: Data Scientist Position, Paris-Saclay Center for Data Science (CDS), France

Balazs mentioned this to me a while ago, here is a position of Data Scientist in his group:
 

Data Scientist Position

Paris-Saclay Center for Data Science (CDS)
The Paris-Saclay CDS for has opened data science positions to reinforce the data science ecosystem and to build and support a data science platform. The data scientist will work on (in order of priority):
  • Data-science support:  The CDS has many collaborations between scientists and engineers across scientific disciplines. The candidate will work on a small number of data-science projects with different scientific partners. The candidate should be interested in a variety of scientific applications of data processing and able to understand an application, to communicate with scientists of a different culture, and to deploy data-processing solutions for their needs. Current projects include: 
    • Galaxy/star classification for preparing the LSST data processing pipeline.
    • Particle tracking for the ATLAS/LHC upgrade.
    • Segmenting and classifying Solar wind time series data.
    • Classifying hospital stays (PMSI coding) with APHP.
    • Laser spectrometry to improve the safety of intravenous drug administration.
    • Improving search in large biological and biomedical data sets.
    • Analyzing sensor data for tracking pollution and its health effects (Polluscope).
    • Classifying crowdsourced ecology data (Spipoll).
  • Software engineering: The objective to enable researchers to do better science thanks to better software tools. It means bringing state of the art data science research software into high quality toolboxes (as scikit-learn). Getting involved with open source development. Assisting data scientists (students, postdoctoral fellows, permanent researchers) to develop their software engineering skills and to get them involved with open source development.
  • Training:  Accompany domain scientists in their data analysis efforts. Accompany data scientists in their methodological research. Designing training sprints and practical material for the courses (cf. based on software carpentry)
Team
Qualifications
  • M.S. / Ph.D. in Computer Science, Statistical Machine Learning

  • Good understanding of the data science workflow. Experience with data challenges is a plus.
  • Strong programming experience with one or more data science languages (Python, R, Matlab)
  • Experience with open source development
 (desired but not required).

The applicant should send a CV, a statement of purpose, and up to three letters of recommendations to cdsupsay@gmail.com.

Position is open now and will be open until it is filled.

Description:
Duration: 18 months
Gross salary per month in €: 2500-2817
 
 
 
 
Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Internship: Signal and image classification with invariant descriptors (scattering transforms), IFPEN Rueil Malmaison, France

Laurent posted an internship position on his site, let me relay it here:
 
Here is a data science internship position at IFP Energies nouvelles (5 months, Spring/Summer 2017), in the field of :

Application and additional details

Description

The field of complex data analysis (data science) is interested in the extraction of suitable indicators used for dimension reduction, data comparison or classification. Initially based on application-dependent, physics-based descriptors or features, novel methods employ more generic and potentially multiscale descriptors, that can be used for machine learning or classification. Examples are to be found in SIFT-like (scale-invariant feature transform) techniques (ORB, SURF), in unsupervised or deep learning. The present internship focuses on the framework of scattering transform (S. Mallat et al.) and the associated classification techniques. It yields signal, image or graph representations with invariance properties relative to data-modifying transformations: translation, rotation, scale… Its performances have been well-studied for classical data (audio signals, image databases, handwritten digit recognition).
This internship aims at dealing with lesser studied data: identification of the closest match to a template image inside a database of underground models, extraction of suitable fingerprints from 1D spectrometric signals from complex chemical compounds for macroscopic chemometric property learning. 
The stake in the first case resides in the different scale and nature of template and model images, the latter being sketch- or cartoon-like versions of the templates. In the second case, signals are composed of a superposition of hundreds of (positive) peaks. Their nature differs from standard information processed by scattering transforms. A focus on one of the proposed applications can be considered, depending on success or difficulties met. 

References


 
 
 
Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Job: Data scientist: deep learning approaches to microbial genomics, Institut Pasteur, Paris

David Bikard posted the following job at the intersection of two explosive fields: 
The Synthetic Biology group is looking for a data scientist to develop deep learning approaches applied to bacterial genomics data. In particular the candidate will work in close collaboration with biologists to analyse high-throughput data generated using CRISPR tools. Applicants should ideally be trained in statistics, machine learning or computer science. No previous knowledge in Biology is required, but you do need to be curious and motivated to learn. Masters, PhDs, Engineers and Postdocs are all welcome to apply. Funding is guaranteed through an ERC starting grant. Applications should be send to david.bikard@pasteur.fr 
Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Job: Postdoc Position in Machine Learning for Health Sensing, MIT

Dina just sent me the following:

Hi Igor,

I hope all is well. I was wondering whether it is possible to list a postdoc opening on your blog. We have a large MIT project the involves medical applications, and we are looking for postdocs with some experience in Machine Learning and/or related fields.  Any help in reaching out to suitable candidates will be highly appreciated. I included the announcement below. 

Thanks a lot and happy Holidays.  -Dina
Sure Dina, here it is:



--

Job Announcement -- Postdoc Position in Machine Learning for Health Sensing

This postdoc position involves working with MIT faculty, MIT students, and medical doctors to develop ML algorithms and software systems to analyze new sensor data and infer disease progression and medication efficacy.  It will build on our prior work on a technology for extracting health metrics by analyzing how human bodies interact with the surrounding wireless signals. The new research will focus on how such new data type can be used in medical applications. In particular, the postdoc will lead a sub-project that deploys this sensor with patients with a particular condition (e.g., Parkinson's), and in collaboration with a medical team and MIT researchers, develop ML algorithms and systems to extract disease markers from the collected data. The research is interdisciplinary; the research outcomes will span ML conferences (ICML, NIPS), wireless system conferences (MobiCom, Sensys), and medical journals.

The position is for a minimum of one year that can be extended to two years. The candidate should have completed (or nearly completed) a PhD in Computer Science, with a focus on machine learning or related fields such as computer vision, speech, or data mining. In particular, expertise in deep neural networks and transfer learning is highly desirable. The candidate should also have expertise in programming and software systems. No medical background is required. Also no background in wireless signals is required. Interested candidates should email Professor Katabi at dk@mit.edu. Please include the word "postdoc2017" in the email title.  
 
 
Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Monday, December 26, 2016

Linearly Convergent Randomized Iterative Methods for Computing the Pseudoinverse -implementation -

Following up on Robert's thesis, here is his latest RandNLA preprint: Linearly Convergent Randomized Iterative Methods for Computing the Pseudoinverse by Robert M. Gower, Peter Richtárik
We develop the first stochastic incremental method for calculating the Moore-Penrose pseudoinverse of a real rectangular matrix. By leveraging three alternative characterizations of pseudoinverse matrices, we design three methods for calculating the pseudoinverse: two general purpose methods and one specialized to symmetric matrices. The two general purpose methods are proven to converge linearly to the pseudoinverse of any given matrix. For calculating the pseudoinverse of full rank matrices we present additional two specialized methods which enjoy faster convergence rate than the general purpose methods. We also indicate how to develop randomized methods for calculating approximate range space projections, a much needed tool in inexact Newton type methods or quadratic solvers when linear constraints are present. Finally, we present numerical experiments of our general purpose methods for calculating pseudoinverses and show that our methods greatly outperform the Newton-Schulz method on large dimensional matrices.
 an implementation is on Github.
 
 
Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Thesis: Sketch and Project: Randomized Iterative Methods for Linear Systems and Inverting Matrices - implementation -

Congratulation Dr. Gower !

On top of his thesis, Robert has a few implementations connected to the thesis, here is an excerpt:

InvRand - download, paper

Inverse Random is suite of randomized methods for inverting positive definite matrices implemented in MATLAB

RandomLinearLab - download, paper

Random Linear Lab is a lab for testing and comparing randomized methods for solving linear systems all implemented in MATLAB

Here is the thesis, with the arxiv abstract, the real abstract and the lay person's abstract !
Sketch and Project: Randomized Iterative Methods for Linear Systems and Inverting Matrices by Robert M. Gower

Probabilistic ideas and tools have recently begun to permeate into several fields where they had traditionally not played a major role, including fields such as numerical linear algebra and optimization. One of the key ways in which these ideas influence these fields is via the development and analysis of randomized algorithms for solving standard and new problems of these fields. Such methods are typically easier to analyze, and often lead to faster and/or more scalable and versatile methods in practice.
This thesis explores the design and analysis of new randomized iterative methods for solving linear systems and inverting matrices. The methods are based on a novel sketch-and-project framework. By sketching we mean, to start with a difficult problem and then randomly generate a simple problem that contains all the solutions of the original problem. After sketching the problem, we calculate the next iterate by projecting our current iterate onto the solution space of the sketched problem.

 Here is the real abstract of the thesis:
Probabilistic ideas and tools have recently begun to permeate into several fields where they had traditionally not played a major role, including fields such as numerical linear algebra and optimization. One of the key ways in which these ideas influence these fields is via the development and analysis of randomized algorithms for solving standard and new problems of these fields. Such methods are typically easier to analyze, and often lead to faster and/or more scalable and versatile methods in practice.
This thesis explores the design and analysis of new randomized iterative methods for solving linear systems and inverting matrices. The methods are based on a novel sketch-and-project framework. By sketching we mean, to start with a difficult problem and then randomly generate a simple problem that contains all the solutions of the original problem. After sketching the problem, we calculate the next iterate by projecting our current iterate onto the solution space of the sketched problem.
The starting point for this thesis is the development of an archetype randomized method for solving linear systems. Our method has six different but equivalent interpretations: sketch-and-project, constrain-and-approximate, random intersect, random linear solve, random update and random fixed point. By varying its two parameters – a positive definite matrix (defining geometry), and a random matrix (sampled in an i.i.d. fashion in each iteration) – we recover a comprehensive array of well known algorithms as special cases, including the randomized Kaczmarz method, randomized Newton method, randomized coordinate descent method and random Gaussian pursuit. We also naturally obtain variants of all these methods using blocks and importance sampling. However, our method allows for a much wider selection of these two parameters, which leads to a number of new specific methods. We prove exponential convergence of the expected norm of the error in a single theorem, from which existing complexity results for known variants can be obtained. However, we also give an exact formula for the evolution of the expected iterates, which allows us to give lower bounds on the convergence rate.
We then extend our problem to that of finding the projection of given vector onto the solution space of a linear system. For this we develop a new randomized iterative algorithm: stochastic dual ascent (SDA). The method is dual in nature, and iteratively solves the dual of the projection problem. The dual problem is a non-strongly concave quadratic maximization problem without constraints. In each iteration of SDA, a dual variable is updated by a carefully chosen point in a subspace spanned by the columns of a random matrix drawn independently from a fixed distribution. The distribution plays the role of a parameter of the method. Our complexity results hold for a wide family of distributions of random matrices, which opens the possibility to fine-tune the stochasticity of the method to particular applications. We prove that primal iterates associated with the dual process converge to the projection exponentially fast in expectation, and give a formula and an insightful lower bound for the convergence rate.
We also prove that the same rate applies to dual function values, primal function values and the duality gap. Unlike traditional iterative methods, SDA converges under virtually no additional assumptions on the system (e.g., rank, diagonal dominance) beyond consistency. In fact, our lower bound improves as the rank of the system matrix drops. By mapping our dual algorithm to a primal process, we uncover that the SDA method is the dual method with respect to the sketch-and-project method from the previous chapter. Thus our new more general convergence results for SDA carry over to the sketch-and-project method and all its specializations (randomized Kaczmarz, randomized coordinate descent...etc). When our method specializes to a known algorithm, we either recover the best known rates, or improve upon them. Finally, we show that the framework can be applied to the distributed average consensus problem to obtain an array of new algorithms. The randomized gossip algorithm arises as a special case.
In the final chapter, we extend our method for solving linear system to inverting matrices, and develop a family of methods with specialized variants that maintain symmetry or positive definiteness of the iterates. All the methods in the family converge globally and exponentially, with explicit rates. In special cases, we obtain stochastic block variants of several quasi-Newton updates, including bad Broyden (BB), good Broyden (GB), Powell-symmetric-Broyden (PSB), Davidon-Fletcher-Powell (DFP) and Broyden-Fletcher-Goldfarb-Shanno (BFGS). Ours are the first stochastic versions of these updates shown to converge to an inverse of a fixed matrix.
Through a dual viewpoint we uncover a fundamental link between quasi-Newton updates and approximate inverse preconditioning. Further, we develop an adaptive variant of the randomized block BFGS (AdaRBFGS), where we modify the distribution underlying the stochasticity of the method throughout the iterative process to achieve faster convergence. By inverting several matrices from varied applications, we demonstrate that AdaRBFGS is highly competitive when compared to the well established Newton-Schulz and approximate preconditioning methods. In particular, on large-scale problems our method outperforms the standard methods by orders of magnitude. The development of efficient methods for estimating the inverse of very large matrices is a much needed tool for preconditioning and variable metric methods in the big data era.
 There is even a Lay Summary
Lay Summary
This thesis explores the design and analysis of methods (algorithms) for solving two common problems: solving linear systems of equations and inverting matrices. Many engineering and quantitative tasks require the solution of one of these two problems. In particular, the need to solve linear systems of equations is ubiquitous in essentially all quantitative areas of human endeavour, including industry and science. Specifically, linear systems are a central problem in numerical linear algebra, and play an important role in computer science, mathematical computing, optimization, signal processing, engineering, numerical analysis, computer vision, machine learning, and many other fields. This thesis proposes new methods for solving large dimensional linear systems and inverting large matrices that use tools and ideas from probability.
The advent of large dimensional linear systems of equations, based on big data sets, poses a challenge. On these large linear systems, the traditional methods for solving linear systems can take an exorbitant amount of time. To address this issue we propose a new class of randomized methods that are capable of quickly obtaining approximate solutions. This thesis lays the foundational work of this new class of randomized methods for solving linear systems and inverting matrices. The main contributions are providing a framework to design and analyze new and existing methods for solving linear systems. In particular, our framework unites many existing methods. For inverting matrices we also provide a framework for designing and analysing methods, but moreover, using this framework we design a highly competitive method for computing an approximate inverse of truly large scale positive definite matrices. Our new method often outperforms previously known methods by several orders of magnitude on large scale matrices 

Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Saturday, December 24, 2016

Saturday Morning Videos: Learning, Algorithm Design and Beyond Worst-Case Analysis workshop, Simons Institute, Berkeley

 
All the videos of the Learning, Algorithm Design and Beyond Worst-Case Analysis workshop at the Simons Institute at Berkeley are available by following the links of each talks.

Learning, Algorithm Design and Beyond Worst-Case Analysis

Nov. 14Nov. 18, 2016

  • Click on the titles of individual talks for abstract, slides and archived video (available approximately one week after the conclusion of the workshop).
Add to My Calendar
Monday, November 14th, 2016
9:00 am – 9:20 am
Coffee and Check-In
9:20 am – 9:30 am
Opening Remarks
9:30 am – 10:10 am
10:10 am – 10:50 am
10:50 am – 11:20 am
Break
11:20 am – 12:00 pm
12:00 pm – 2:00 pm
Lunch
2:00 pm – 2:40 pm
2:40 pm – 3:20 pm
3:20 pm – 3:50 pm
Break
3:50 pm – 4:30 pm
4:40 pm – 5:00 pm
Impromptu Talks Session
5:00 pm – 6:00 pm
Reception
Tuesday, November 15th, 2016
9:00 am – 9:30 am
Coffee and Check-In
9:30 am – 10:10 am
10:10 am – 10:50 am
10:50 am – 11:20 am
Break
11:20 am – 12:00 pm
12:00 pm – 2:00 pm
Lunch
2:00 pm – 2:40 pm
2:40 pm – 3:20 pm
3:20 pm – 3:50 pm
Break
3:50 pm – 4:30 pm
4:40 pm – 5:00 pm
Impromptu Talks Session
Wednesday, November 16th, 2016
9:00 am – 9:30 am
Coffee and Check-In
9:30 am – 10:10 am
10:10 am – 10:50 am
10:50 am – 11:20 am
Break
11:20 am – 12:00 pm
12:00 pm – 12:40 pm
12:40 pm – 2:00 pm
Lunch
2:00 pm – 3:20 pm
Breakout Groups
3:20 pm – 4:00 pm
4:00 pm – 5:00 pm
Discussion
Thursday, November 17th, 2016
9:00 am – 9:30 am
Coffee and Check-In
9:30 am – 10:10 am
10:10 am – 10:50 am
10:50 am – 11:20 am
Break
11:20 am – 12:00 pm
12:00 pm – 2:00 pm
Lunch
2:00 pm – 2:40 pm
2:40 pm – 3:20 pm
3:20 pm – 3:50 pm
Break
3:50 pm – 4:30 pm
4:30 pm – 5:10 pm
5:10 pm – 5:30 pm
Impromptu Talks Session
Friday, November 18th, 2016
9:00 am – 9:30 am
Coffee and Check-In
9:30 am – 10:10 am
10:10 am – 10:50 am
10:50 am – 11:20 am
Break
11:20 am – 12:00 pm
12:00 pm – 2:00 pm
Lunch
2:00 pm – 2:40 pm
2:40 pm – 3:20 pm
3:20 pm – 3:50 pm
Break
3:50 pm – 4:30 pm
4:40 pm – 5:00 pm
Impromptu Talks Session
 
 
 
 
 
 
Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Printfriendly