Advanced Fault Injection
Supports Byzantine failures, network partitioning, cascading failures, and hardware-level faults such as bit flips and NIC drops to test system resilience in highly volatile environments.
RESILIENCE / FAULTSDDLSim-Lab is a distributed simulation environment designed for experimenting with deep learning systems, network conditions, failures, scheduling, edge environments, and system resilience.
The environment brings together configurable network behavior, failure scenarios, compute resources, and distributed workloads so researchers and developers can explore system behavior under controlled conditions.
DDLSim-Lab welcomes researchers, research groups, laboratories, developers, students, engineers, and independent contributors interested in distributed deep learning, simulation, networking, systems, or AI infrastructure.
You can use the project to explore an idea, reproduce an experiment, build an extension, report an issue, propose a research direction, or contribute improvements to the codebase.
Whether you are an individual researcher or part of a university, laboratory, company, or open-source community, you are welcome to participate.
Supports Byzantine failures, network partitioning, cascading failures, and hardware-level faults such as bit flips and NIC drops to test system resilience in highly volatile environments.
RESILIENCE / FAULTSTrain reinforcement learning agents to dynamically allocate compute resources and bandwidth based on workload characteristics and real-time network constraints.
RL / SCHEDULINGModel heterogeneous edge-cloud hierarchies with constrained devices, intermittent connectivity, and strict latency bounds to represent modern IoT setups and federated learning architectures.
EDGE / CLOUD / IOTIncludes dedicated modules for simulating adversarial network attacks, intrusion detection systems, gradient leakage, and secure enclave training environments across the cluster.
SECURITY / AIThe platform can support different research questions around distributed learning, systems, networking, reliability, and infrastructure.
Study distributed deep learning behavior across heterogeneous nodes and changing system conditions.
Investigate the effects of latency, packet loss, bandwidth constraints, and network instability.
Explore how distributed systems respond to failures and degraded infrastructure.
Experiment with constrained devices, distributed data, and heterogeneous edge-cloud environments.
Investigate strategies for allocating compute, bandwidth, and other distributed resources.
Explore security scenarios affecting distributed AI workloads and their surrounding infrastructure.
DDLSim-Lab is open to contributions from people working across research, engineering, academia, and open source. Contributions can range from a small bug fix to a new experiment, module, configuration, documentation improvement, or research direction.