Programing AI Accelerators for Scientific Computing

Programing AI Accelerators for Scientific Computing

There is a lot more to modern computer hardware than just CPUs and GPUs. How can we use all the new technology effectively?

A new trend in processor architecture design is rise of dedicated accelerator hardware for machine learning and HPC applications such as the Cerebras Wafer Scale Engine, Tenstorrent Blackhole, and the NextSilicon Maverick. These processors have evolved from the experimental state into market-ready products, and they have the potential to constitute the next major architectural shift after GPUs saw widespread adoption more than a decade ago.

The most important aspect of these tile-centric processors is the reliance on SRAM as user-controlled memory, allowing for extremely fast memory operations, which have become a bottleneck in traditional architectures. However, their distributed nature of this memory makes it necessary to redesign algorithms. In many cases, algorithms for such accelerators resemble scalable algorithms for supercomputers.

In this thesis we will work with the new hardware and the programming techniques that are required to unlock their potential. We will work on implementing basic algorithms and study how performance can be achieved and analyzed on such devices, and how AI workloads can be accelerated on such systems.

Goal

The goal of the thesis is to get a deeper understanding of the performance characteristics of modern AI hardware, and to report those results.

Learning outcome

  • Performance analysis
  • Low level programing
  • Code optimization

Qualifications

  • C programing
  • Computer architecture

Supervisors

  • Johannes Langguth

Associated contact