OPEN-SOURCE RESEARCH ENVIRONMENT

Distributed deep learning, studied at system scale.

DDLSim-Lab provides a reproducible environment for studying distributed deep learning across heterogeneous infrastructure, network conditions, scheduling strategies, and failure scenarios.

Project lead: Kaitlyn Brishae Truby  ·  Active research since 2025
Open Source Research model
Bare Metal Infrastructure focus
Edge → Cloud Experiment scope
Reproducible Experimentation
01 / Project

A research environment built for distributed systems.

DDLSim-Lab is an open-source research project focused on understanding how distributed deep learning systems behave under realistic infrastructure constraints.

The platform is designed to support experimentation with large-scale, heterogeneous, and failure-prone environments.

Researchers can investigate algorithms, scheduling policies, communication behavior, resource allocation and fault-tolerance techniques.

Systems-level experimentation

Study the interaction between distributed workloads and the infrastructure beneath them.

$ workload → scheduler → network → nodes → telemetry
02 / Research modules

What the lab can investigate

The environment brings several system dimensions together so experiments can examine their interaction.

Distributed Training

Distributed deep learning workloads across multiple compute nodes.

Network Behavior

Latency, jitter, packet loss, bandwidth and communication behavior.

Scheduling

Placement, resource allocation and adaptive scheduling strategies.

Performance

Scaling behavior and system performance under changing conditions.

Fault Tolerance

Failure scenarios, recovery and resilient distributed behavior.

Edge–Cloud Systems

Workloads spanning edge resources and cloud infrastructure.

03 / Research workflow

From experiment design to observation.

A simple research flow keeps experiments understandable and reproducible.

01

Define workload

Specify the distributed workload, resources and experimental conditions.

02

Configure system

Configure nodes, networking, scheduling and infrastructure behavior.

03

Run experiment

Execute the workload while observing system and network behavior.

04

Analyze results

Compare measurements and evaluate the behavior of the system.

04 / System architecture

A systems view of the experiment.

The architecture connects workload, orchestration, network behavior, compute nodes and experiment telemetry.

Workload Deep learning
Scheduler Placement
Network Communication
Compute Nodes Execution
05 / Infrastructure

Bare-metal experimentation

When infrastructure is part of the experiment.

DDLSim-Lab is designed to support bare-metal deployments when experiments require direct access to networking, compute and kernel-level behavior.

  • Reduced virtualization overhead for performance-sensitive measurements.
  • Direct access to network interfaces for realistic communication experiments.
  • Support for technologies such as SR-IOV and RDMA where required.
06 / Open research

Built to be inspected, reproduced and extended.

DDLSim-Lab is developed as an open-source research project. The implementation can evolve with new experiments, research questions and contributions.