The deep learning revolution is in full swing, casting powerful predictive analysis to a slew of industries with unique datasets and accelerating innovation, but not without acknowledgement of the increasing utility of supportive tools and platforms. Software advances manifested in the market with new frameworks to train neural networks and models, like Caffe, Pytorch and Tensorflow. But now the spotlight is on hardware. Dedicated computing deep learning chips are beginning to enter the market, for cloud processing and edge environments, like Graphcore, Horizon.ai, Wave Computing and Cerebras System, competing with the giants like Nvidia, Intel, Google, Qualcomm, Xilinx, AMD, and CEVA, all producing impressive results, yet all within an envelope of tradeoffs. Cloud dedicated deep learning chips are designed to achieve processing high throughput, while deep learning chips for the edge address factors like power efficiency, latency, bandwidth and optimized designs to fit into edge devices. Israeli startup Hailo is doubling down a fundamental market inefficiency with their new deep learning chip that’s designed from a new processing architecture and enables running efficient deep learning computing on the edge.
[caption id="attachment_54633" align="alignnone" width="1184"] Hailo's team, Tel Aviv office. Credit: Hailo[/caption]
Hailo, founded in 2017, is a fabless semiconductor company developing deep learning chips designed to deliver data-center performance on edge devices. Their technology allows their chip to distribute compute problems across the chip’s sections, enabling them to consume extra low power for deep learning computations.
[caption id="attachment_54635" align="alignnone" width="1248"] Hailo-8. Credit: Hailo.[/caption]
Their first product, the Hailo-8, launched in May 2019, is arguably the world’s fastest and most efficient deep learning processor for edge devices. It registers a maximum capacity of 26 TOPS (Tera Operations Per Second), with measured 3 TOPS per watt efficiency. And it promises 20 times better efficiency than nvidia’s Xavier AGX, all inside a significantly smaller form factor. They measured power consumption of only 1.7 watts by running a ResNet-50 on low resolution video (224 x 224, 8-bit precision with a batch size of 1) at 672 fps.
