Thursday, April 13, 2017

In-Datacenter Performance Analysis of a Tensor Processing Unit​ (TM)

Norm Jouppi mentioned it in a Google blog entry a few days. He and his team have released some information about the TPU. This is very interesting because of how algorithms are changing hardware. Sure enough we know that NVIDIA has shifted its architecture to follow the successful Deep Neural networks development (and continues to do so) but when Google announced the TPU last year one could only wonder what they would be doing better than a chip maker.    wrote about this question in  Does Google’s TPU Investment Make Sense Going Forward? but I believe he misses the point entirely.

That point was made abundantly clear in a Wired article:
It’s not used to train the neural network beforehand. But as Jouppi explains, even that still saves the company quite a bit. It didn’t have to build, say, an extra 15 data centers
The reason Google did not wait for NVIDIA's architecture to change or for Moore's law to kick in (that is using OPM to do the work) is mostly because that hardware technology effort spare them a few bucks. It all comes down to the fact that we are all collectively not fast enough to make sense of data.

Here is Norm et al's paper:  In-Datacenter Performance Analysis of a Tensor Processing Unit​ (TM) by Norman P. Jouppi, Cliff Young, Nishant Patil, David Patterson, Gaurav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Boden, Al Borchers, Rick Boyle, Pierre-luc Cantin, Clifford Chao, Chris Clark, Jeremy Coriell, Mike Daley, Matt Dau, Jeffrey Dean, Ben Gelb, Tara Vazir Ghaemmaghami, Rajendra Gottipati, William Gulland, Robert Hagmann, C. Richard Ho, Doug Hogberg, John Hu, Robert Hundt, Dan Hurt, Julian Ibarz, Aaron Jaffey, Alek Jaworski, Alexander Kaplan, Harshit Khaitan, Andy Koch, Naveen Kumar, Steve Lacy, James Laudon, James Law, Diemthu Le, Chris Leary, Zhuyuan Liu, Kyle Lucke, Alan Lundin, Gordon MacKean, Adriana Maggiore, Maire Mahony, Kieran Miller, Rahul Nagarajan, Ravi Narayanaswami, Ray Ni, Kathy Nix, Thomas Norrie, Mark Omernick, Narayana Penukonda, Andy Phelps, Jonathan Ross, Matt Ross, Amir Salek, Emad Samadiani, Chris Severn, Gregory Sizikov, Matthew Snelham, Jed Souter, Dan Steinberg, Andy Swing, Mercedes Tan, Gregory Thorson, Bo Tian, Horia Toma, Erick Tuttle, Vijay Vasudevan, Richard Walter, Walter Wang, Eric Wilcox, and Doe Hyun Yoon

Many architects believe that major improvements in cost-energy-performance must now come from domain-specific hardware. This paper evaluates a custom ASIC—called a ​Tensor Pro​cessing Unit (TPU)— deployed in datacenters since 2015 that accelerates the inference phase of neural networks (NN). The heart of the TPU is a 65,536 8-bit MAC matrix multiply unit that offers a peak throughput of 92 TeraOps/second (TOPS) and a large (28 MiB) software-managed on-chip memory. The TPU’s deterministic execution model is a better match to the 99th-percentile response-time requirement of our NN applications than are the time-varying optimizations of CPUs and GPUs (caches, out-of-order execution, multithreading, multiprocessing, prefetching, ...) that help average throughput more than guaranteed latency. The lack of such features helps explain why, despite having myriad MACs and a big memory, the TPU is relatively small and low power. We compare the TPU to a server-class Intel Haswell CPU and an Nvidia K80 GPU, which are contemporaries deployed in the same datacenters. Our workload, written in the high-level TensorFlow framework, uses production NN applications (MLPs, CNNs, and LSTMs) that represent 95% of our datacenters’ NN inference demand. Despite low utilization for some applications, the TPU is on average about 15X - 30X faster than its contemporary GPU or CPU, with TOPS/Watt about 30X - 80X higher. Moreover, using the GPU’s GDDR5 memory in the TPU would triple achieved TOPS and raise TOPS/Watt to nearly 70X the GPU and 200X the CPU.







Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Wednesday, April 12, 2017

Two jobs, Research Scientists, NVIDIA

 This is new but since I did post a similar job opening for LightOn  in the realm of Machine Learning and Hardware, NVIDIA falls indeed in the same category:

Hi Igor,


Here are two great NVIDIA research roles, if you are interested in posting:

https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Research-Scientist--Machine-Learning-_JR1905891?shared_id=e998709d-4743-42e4-b2db-15a939ddd52a

https://nvidia.wd5.myworkdayjobs.com/NVIDIAExternalCareerSite/job/US-CA-Santa-Clara/Research-Scientist--Computer-Vision_JR1905140?shared_id=fe68f03b-f0d0-448a-829a-ba363b71f531

Let me know!

Best Regards,

Anita

Anita Rexinger| Talent Acquisition | NVIDIA Corporation
 
 
 
 
 
 
Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Highly Technical Reference Page: The Incredible PyTorch


Ritchie Ng put together The Incredible PyTorch, a curated list of tutorials, papers, projects, communities and more relating to PyTorch.It is here: https://github.com/ritchieng/the-incredible-pytorch
 
The reference page will be added to the list of Highly Technical Reference Pages


h/t Desernaut




Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Thursday, April 06, 2017

CSVideoNet: A Real-time End-to-end Learning Framework for High-frame-rate Video Compressive Sensing - implementation -

The Great Convergence continues: Compressive Sensing allowed the encoder to be non iterative and now a set of neural networks (CNNs, LSTMs) make the decoding also non iterative. Woohoo !


Fengbo who just received an interesting NSF award also sent me the following:



Dear Igor, 
Hope all is well.
Recently, we proposed a deep learning based encoding/decoding framework (CSVideoNet) for high-frame-rate video compressive sensing. We believe you and your blog’s readers would be interested in this research. The preview of the manuscript can be found here: https://arxiv.org/abs/1612.05203

CSVideoNet: A Real-time End-to-end Learning Framework for High-frame-rate Video Compressive Sensing by Kai Xu, Fengbo Ren
This paper addresses the real-time encoding-decoding problem for high-frame-rate video compressive sensing (CS). Unlike prior works that perform reconstruction using iterative optimization-based approaches, we propose a non-iterative model, named "CSVideoNet". CSVideoNet directly learns the inverse mapping of CS and reconstructs the original input in a single forward propagation. To overcome the limitations of existing CS cameras, we propose a multi-rate CNN and a synthesizing RNN to improve the trade-off between compression ratio (CR) and spatial-temporal resolution of the reconstructed videos. The experiment results demonstrate that CSVideoNet significantly outperforms the state-of-the-art approaches. With no pre/post-processing, we achieve 25dB PSNR recovery quality at 100x CR, with a frame rate of 125 fps on a Titan X GPU. Due to the feedforward and high-data-concurrency natures of CSVideoNet, it can take advantage of GPU acceleration to achieve three orders of magnitude speed-up over conventional iterative-based approaches. We share the source code at this https URL
Thanks,
================================
Fengbo​ Ren
Director, Parallel Systems and Computing Laboratory (PSCLab)
​Assistant Professor, School of Computing, Informatics, and Decision Systems Engineering (CIDSE)
Arizona State University (ASU)
========================​​=====​​===

Thanks Fengbo !


Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Tuesday, April 04, 2017

Improved Training of Wasserstein GANs - implementation -


Thanks Alex for the heads-up !


Generative Adversarial Networks (GANs) are powerful generative models, but suffer from training instability. The recently proposed Wasserstein GAN (WGAN) makes significant progress toward stable training of GANs, but can still generate low-quality samples or fail to converge in some settings. We find that these training failures are often due to the use of weight clipping in WGAN to enforce a Lipschitz constraint on the critic, which can lead to pathological behavior. We propose an alternative method for enforcing the Lipschitz constraint: instead of clipping weights, penalize the norm of the gradient of the critic with respect to its input. Our proposed method converges faster and generates higher-quality samples than WGAN with weight clipping. Finally, our method enables very stable GAN training: for the first time, we can train a wide variety of GAN architectures with almost no hyperparameter tuning, including 101-layer ResNets and language models over discrete data.


Implementation is here: https://github.com/igul222/improved_wgan_training


Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !

It's all about convolutions: YodaNN, TrueNorth computing, FPGAs and a comparison between HOG and CNNs on Hardware

After yesterday's Compressive Sensing Hardware, let us look at the recent hardware needed to speed up Deep Learning algorithms where convolutions are a bottelneck. The last paper is also a welcome addition to people wondering about the link between architectures based on DL and those based on computer vision features. Enjoy !




YodaNN: An Architecture for Ultra-Low Power Binary-Weight CNN Acceleration by Renzo Andri, Lukas Cavigelli, Davide Rossi, Luca Benini


Convolutional neural networks (CNNs) have revolutionized the world of computer vision over the last few years, pushing image classification beyond human accuracy. The computational effort of today's CNNs requires power-hungry parallel processors or GP-GPUs. Recent developments in CNN accelerators for system-on-chip integration have reduced energy consumption significantly. Unfortunately, even these highly optimized devices are above the power envelope imposed by mobile and deeply embedded applications and face hard limitations caused by CNN weight I/O and storage. This prevents the adoption of CNNs in future ultra-low power Internet of Things end-nodes for near-sensor analytics. Recent algorithmic and theoretical advancements enable competitive classification accuracy even when limiting CNNs to binary (+1/-1) weights during training. These new findings bring major optimization opportunities in the arithmetic core by removing the need for expensive multiplications, as well as reducing I/O bandwidth and storage. In this work, we present an accelerator optimized for binary-weight CNNs that achieves 1510 GOp/s at 1.2 V on a core area of only 1.33 MGE (Million Gate Equivalent) or 0.19 mm
2

and with a power dissipation of 895 {\mu}W in UMC 65 nm technology at 0.6 V. Our accelerator significantly outperforms the state-of-the-art in terms of energy and area efficiency achieving 61.2 TOp/s/W@0.6 V and 1135 GOp/s/MGE@1.2 V, respectively.




  Deep networks are now able to achieve human-level performance on a broad spectrum of recognition tasks. Independently, neuromorphic computing has now demonstrated unprecedented energy-efficiency through a new chip architecture based on spiking neurons, low precision synapses, and a scalable communication network. Here, we demonstrate that neuromorphic computing, despite its novel architectural primitives, can implement deep convolution networks that (i) approach state-of-the-art classification accuracy across eight standard datasets encompassing vision and speech, (ii) perform inference while preserving the hardware’s underlying energy-efficiency and high throughput, running on the aforementioned datasets at between 1,200 and 2,600 frames/s and using between 25 and 275 mW (effectively  superior to 6,000 frames/s per Watt), and (iii) can be specified and trained using backpropagation with the same ease-of-use as contemporary deep learning. This approach allows the algorithmic power of deep learning to be merged with the efficiency of neuromorphic processors, bringing the promise of embedded, intelligent, brain-inspired computing one step closer.


FPGA-based hardware accelerators for convolutional neural networks (CNNs) have obtained great attentions due to their higher energy efficiency than GPUs. However, it is challenging for FPGA-based solutions to achieve a higher throughput than GPU counterparts. In this paper, we demonstrate that FPGA acceleration can be a superior solution in terms of both throughput and energy efficiency when a CNN is trained with binary constraints on weights and activations. Specifically, we propose an optimized accelerator architecture tailored for bitwise convolution and normalization that features massive spatial parallelism with deep pipelines stages. Experiment results show that the proposed architecture is 8.3x faster and 75x more energy-efficient than a Titan X GPU for processing online individual requests (in small batch size). For processing static data (in large batch size), the proposed solution is on a par with a Titan X GPU in terms of throughput while delivering 9.5x higher energy efficiency.

Towards Closing the Energy Gap Between HOG and CNN Features for Embedded Vision by Amr Suleiman, Yu-Hsin Chen, Joel Emer, Vivienne Sze

Computer vision enables a wide range of applications in robotics/drones, self-driving cars, smart Internet of Things, and portable/wearable electronics. For many of these applications, local embedded processing is preferred due to privacy and/or latency concerns. Accordingly, energy-efficient embedded vision hardware delivering real-time and robust performance is crucial. While deep learning is gaining popularity in several computer vision algorithms, a significant energy consumption difference exists compared to traditional hand-crafted approaches. In this paper, we provide an in-depth analysis of the computation, energy and accuracy trade-offs between learned features such as deep Convolutional Neural Networks (CNN) and hand-crafted features such as Histogram of Oriented Gradients (HOG). This analysis is supported by measurements from two chips that implement these algorithms. Our goal is to understand the source of the energy discrepancy between the two approaches and to provide insight about the potential areas where CNNs can be improved and eventually approach the energy-efficiency of HOG while maintaining its outstanding performance accuracy.

 
Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Monday, April 03, 2017

Lensless Imaging with Compressive Ultrafast Sensing

Here is a new compressive architecture that uses CMOS SPADs.



Lensless Imaging with Compressive Ultrafast Sensing by Guy Satat, Matthew Tancik, Ramesh Raskar
Lensless imaging is an important and challenging problem. One notable solution to lensless imaging is a single pixel camera which benefits from ideas central to compressive sampling. However, traditional single pixel cameras require many illumination patterns which result in a long acquisition process. Here we present a method for lensless imaging based on compressive ultrafast sensing. Each sensor acquisition is encoded with a different illumination pattern and produces a time series where time is a function of the photon's origin in the scene. Currently available hardware with picosecond time resolution enables time tagging photons as they arrive to an omnidirectional sensor. This allows lensless imaging with significantly fewer patterns compared to regular single pixel imaging. To that end, we develop a framework for designing lensless imaging systems that use ultrafast detectors. We provide an algorithm for ideal sensor placement and an algorithm for optimized active illumination patterns. We show that efficient lensless imaging is possible with ultrafast measurement and compressive sensing. This paves the way for novel imaging architectures and remote sensing in extreme situations where imaging with a lens is not possible.

The page for the projet is here: http://web.media.mit.edu/~guysatat/singlepixel/


Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !

Nuit Blanche in Review (March 2017)

So this past month, since the last Nuit Blanche in Review (February 2017), we asked ourselves "How can you tell the world is changing right before your eyes ?", passed six million page views on Nuit Blanche, advertized a job in Machine Learning at LightOn (as well as many others see the job section), saw one potential impact of AI in the future (Warming Up ) and featured the first Compressed Sensing preprint that uses Generative Models. We also featured several announcements (Summer schools), had two Paris Machine Learning meetups, and had several videos. Enjoy ! 




Sunday Morning Insight

Implementations

The blogs;

In-depth:

Job
Meetup
Announcements:

Videos:

Other:


Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Saturday, April 01, 2017

Saturday Morning Video: Deep Learning And The Future Of AI, Yann LeCun at Tsinghua University

Yann  mentioned that the video  and slides of his lecture this week  at Tsinghua University is available. Here are the slides


 
Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Saturday Morning Videos: Representation Learning Workshop at Simons Institute, Berkeley (March 27th-31st, 2017)

A photo of New Zealand taken this week. 

 ICLR is in three weeks, so what about dwelving into representation learning. As it turns out, this week, the Representation Learning Workshop took place at the Simons Institute for the Theory of Computation at UC Berkeley. Here are all the videos (an impressive feat given it happened this week). Thank you to the organizers Sham Kakade , Sanjeev Arora , Kristen Grauman, Ruslan Salakhutdinov, Noah Smith  for the line-up. 
Here is the abstract:
This workshop will focus on dramatic advances in representation and learning taking place in natural language processing, speech and vision.  For instance, deep learning can be thought of as a method that combines the tasks of finding a classifier (which we can think of as the top layer of the deep net) with the task of learning a representation (namely, the representation computed at the last-but-one layer).

Developing a theory for such empirical work is an exciting quest, especially since the empirical work draws upon non-convex optimization.  The workshop will draw a mix of theorists and practitioners, and the following is a list of sample issues that will be discussed:  (a) Which models for representation make more sense than
others, and why? (In other words, what patterns in data are they capturing, and how are those patterns useful?)  (b) What is an analog of generalization theory for representation learning?  Can it lead to a theory of transfer learning to new distributions of inputs?  (c) How can we design algorithms for representation learning with provable guarantees?  What progress has already been made, and what lessons can we draw from it?  (d) How can we learn representations that combine probabilities and logic?
and the videos (Russ's slides are also added per his tweet) . Enjoy !
Monday, March 27th, 2017
9:00 am – 9:20 am
Coffee and Check-In
9:20 am – 9:30 am
Opening Remarks
9:30 am – 10:10 am
10:10 am – 10:50 am
10:50 am – 11:20 am
Break
11:20 am – 12:20 pm
12:20 pm – 2:15 pm
Lunch
2:15 pm – 3:15 pm
3:15 pm – 3:45 pm
Break
3:45 pm – 4:25 pm
4:30 pm – 4:45 pm
4:45 pm – 5:00 pm
5:00 pm – 6:00 pm
Reception
Tuesday, March 28th, 2017
9:00 am – 9:30 am
Coffee and Check-In
9:30 am – 10:10 am
10:10 am – 10:50 am
10:50 am – 11:20 am
Break
11:20 am – 12:20 pm
12:20 pm – 2:15 pm
Lunch
2:15 pm – 2:55 pm
2:55 pm – 3:35 pm
3:35 pm – 4:10 pm
Break
4:10 pm – 5:00 pm
Panel
Wednesday, March 29th, 2017
9:00 am – 9:30 am
Coffee and Check-In
9:30 am – 10:10 am
10:10 am – 10:50 am
10:50 am – 11:20 am
Break
11:20 am – 12:20 pm
12:20 pm – 2:15 pm
Lunch
2:15 pm – 2:55 pm
2:55 pm – 3:35 pm
3:35 pm – 4:10 pm
Break
4:10 pm – 4:30 pm
4:35 pm – 4:55 pm
Thursday, March 30th, 2017
9:00 am – 9:30 am
Coffee and Check-In
9:30 am – 10:10 am
10:10 am – 10:50 am
10:50 am – 11:20 am
Break
11:20 am – 12:20 pm
12:20 pm – 2:15 pm
Lunch
2:15 pm – 2:55 pm
2:55 pm – 3:35 pm
3:35 pm – 4:10 pm
Break
4:10 pm – 5:00 pm
Panel
Friday, March 31st, 2017
9:00 am – 9:30 am
Coffee and Check-In
9:30 am – 10:10 am
10:10 am – 10:50 am
10:50 am – 11:20 am
Break
11:20 am – 12:20 pm
12:20 pm – 2:15 pm
Lunch
2:15 pm – 2:55 pm
2:55 pm – 3:35 pm
3:35 pm – 4:10 pm
Break
4:10 pm – 4:30 pm





Join the CompressiveSensing subreddit or the Google+ Community or the Facebook page and post there !
Liked this entry ? subscribe to Nuit Blanche's feed, there's more where that came from. You can also subscribe to Nuit Blanche by Email, explore the Big Picture in Compressive Sensing or the Matrix Factorization Jungle and join the conversations on compressive sensing, advanced matrix factorization and calibration issues on Linkedin.

Printfriendly