Explore recent technical work covering
distributed-system failures, intelligent
scheduling and realistic edge AI simulation.
3 research posts
FEATURED RESEARCH NOTE
Advances in Fault Injection Techniques
for Distributed AI Systems
DDLSim-Lab's fault injection capabilities
have been expanded to support more
sophisticated distributed-system failure
models, including Byzantine behavior,
network partitions and cascading failures.
March 20, 2026
•
Technical research
Fault Injection
AI Resilience
Mar 20, 2026
•
Fault Injection
AI Resilience
Advances in Fault Injection Techniques
for Distributed AI Systems
In our latest research, we've enhanced
DDLSim-Lab's fault injection capabilities
to support more sophisticated failure
models. Beyond traditional latency and
packet loss, we now simulate Byzantine
failures, network partitions and cascading
failures.
These advanced fault models enable
researchers to test distributed training
algorithms under adverse conditions.
Mar 15, 2026
•
AI Scheduling
Edge Computing
AI-Driven Scheduling for
Heterogeneous Edge-Cloud Training
We've implemented a reinforcement
learning framework within DDLSim-Lab
that dynamically allocates computational
and network resources across heterogeneous
edge-cloud infrastructures.
The system treats each node as an agent
in a multi-agent reinforcement learning
environment.
Mar 10, 2026
•
Edge AI
Federated Learning
Edge AI Simulation: Enabling Realistic
Federated Learning Experiments
DDLSim-Lab now supports realistic
simulation of edge AI scenarios where
constrained devices with intermittent
connectivity participate in federated
learning processes.
Researchers can define custom device
profiles representing real-world IoT
deployments.
No research notes found
Try another search term or research category.
02 / Research topics
Areas we write about.
Fault Tolerance
Failure models, resilience,
recovery behavior and distributed
system reliability.
AI Scheduling
Intelligent resource allocation,
reinforcement learning and
heterogeneous workloads.
Distributed Networking
Latency, packet loss, network
behavior and communication
bottlenecks.
Edge AI
Resource-constrained devices,
edge-cloud environments and
distributed inference.
Federated Learning
Device heterogeneity, aggregation,
stragglers and adaptive training.
Experimental Analysis
Convergence, accuracy, recovery
time and system-level measurements.
03 / Explore further
Go deeper into the project.
Explore the research direction,
technical documentation and source
code behind DDLSim-Lab.