Repository logo
  • Log In
    Log in via Symplectic to deposit your publication(s).
Repository logo
  • Communities & Collections
  • Research Outputs
  • Statistics
  • Log In
    Log in via Symplectic to deposit your publication(s).
  1. Home
  2. Faculty of Engineering
  3. Faculty of Engineering
  4. Making State Explicit for Imperative Big Data Processing
 
  • Details
Making State Explicit for Imperative Big Data Processing
File(s)
sdg-atc14-final.pdf (517.4 KB)
Published version
Author(s)
Pietzuch, PR
Fernandez, RC
Migliavacca, M
Kalyvianaki, E
Type
Conference Paper
Abstract
Data scientists often implement machine learning algorithms in imperative languages such as Java, Matlab and R. Yet such implementations fail to achieve the performance and scalability of specialised data-parallel processing frameworks. Our goal is to execute imperative Java programs in a data-parallel fashion with high
throughput and low latency. This raises two challenges: how to support the arbitrary mutable state of Java programs without compromising scalability, and how to re
cover that state after failure with low overhead. Our idea is to infer the dataflow and the types of state accesses from a Java program and use this information
to generate a stateful dataflow graph (SDG) . By explicitly separating data
from mutablestate, SDGs have specific features to enable this translation: to ensure scalability, distributed state can be partitioned across nodes if computation can occur entirely in parallel; if this is not possible, partial state gives nodes local instances for independent computation, which are reconciled according
to application semantics. For fault tolerance, large inmemory state is checkpointed asynchronously without global coordination. We show that the performance of
SDGs for several imperative online applications matches
that of existing data-parallel processing frameworks.
Date Issued
2014-06-01
Date Acceptance
2014-04-01
Citation
Proceedings of USENIX ATC ’14: 2014 USENIX Annual Technical Conference, 2014, pp.49-60
URI
http://hdl.handle.net/10044/1/62212
ISBN
978-1-931971-10-2
Publisher
USENIX
Start Page
49
End Page
60
Journal / Book Title
Proceedings of USENIX ATC ’14: 2014 USENIX Annual Technical Conference
Copyright Statement
© 2014 The Author(s). This is an open access article.
Sponsor
Engineering & Physical Science Research Council (EPSRC)
Identifier
https://lsds.doc.ic.ac.uk/sites/default/files/sdg-atc14-final.pdf
Grant Number
EP/F035217/1
Source
USENIX Annual Technical Conference (USENIX ATC)
Publication Status
Published
Start Date
2014-06-19
Finish Date
2014-06-20
Coverage Spatial
Philadelphia, PA, USA
Date Publish Online
2014-06-01
About
Spiral Depositing with Spiral Publishing with Spiral Symplectic
Contact us
Open access team Report an issue
Other Services
Scholarly Communications Library Services
logo

Imperial College London

South Kensington Campus

London SW7 2AZ, UK

tel: +44 (0)20 7589 5111

Accessibility Modern slavery statement Cookie Policy

Built with DSpace-CRIS software - Extension maintained and optimized by 4Science

  • Cookie settings
  • Privacy policy
  • End User Agreement
  • Send Feedback