Improving the performance of dataflow systems for deep neural network training

Watcharapichat, Pijika

doi:https://doi.org/10.25560/57955

File(s)

Watcharapichat-P-2018-PhD-Thesis.pdf (1.78 MB)

Thesis

Author(s)

Watcharapichat, Pijika

Type

Thesis or dissertation

Abstract

Deep neural networks (DNNs) have led to significant advancements in machine learning.
With deep structure and flexible model parameterisation, they exhibit state-of-the-art accuracies for many complex tasks e.g. image recognition. To achieve this, models are trained iteratively over large datasets. This process involves expensive matrix operations, making it time-consuming to obtain converged models. To accelerate training, dataflow systems parallelise computation. A scalable approach is to use parameter server framework: it has workers that train model replicas in parallel and parameter servers that synchronise the replicas to ensure the convergence.

With distributed DNN systems, there are three challenges that determine the training completion time. In this thesis, we propose practical and effective techniques to address each of these challenges.

Since frequent model synchronisation results in high network utilisation, the parameter server approach can suffer from network bottlenecks, thus requiring decisions on resource allocation. Our idea is to use all available network bandwidth and synchronise subject to the available bandwidth. We present Ako, a DNN system that uses partial gradient exchange for synchronising replicas in a peer-to-peer fashion. We show that our technique exhibits a 25% lower convergence time than a hand-tuned parameter-server deployments.

For a long training, the compute efficiency of worker nodes is important. We argue that processing hardware should be fully utilised for the best speed-up. The key observation is it is possible to overlap the execution of several matrix operations with other workloads. We describe Crossbow, a GPU-based system that maximises hardware utilisation. By using a multi-streaming scheduler, multiple models are trained in parallel on GPU and achieve a 2.3x speed-up compared to a state-of-the-art system.

The choice of model configuration for replicas also directly determines convergence quality. Dataflow systems are used for exploring the promising configurations but provide little support for efficient exploratory workflows. We present Meta-dataflow (MDF), a dataflow model that expresses complex workflows. By taking into account all configurations as a unified workflow, MDFs efficiently reduce time spent on configuration exploration.

Version

Open Access

Date Issued

2017-09

Date Awarded

2018-02

URI

http://hdl.handle.net/10044/1/57955

DOI

https://doi.org/10.25560/57955

Advisor

Pietzuch, Peter

Sponsor

Imperial College London

Publisher Department

Computing

Publisher Institution

Imperial College London

Qualification Level

Doctoral

Qualification Name

Doctor of Philosophy (PhD)