casRL: A Compiler and Runtime for Composable Reinforcement Learning Systems
Murtaza Rangwala and R. K. Williams
Manuscript in preparation; targeting MLSys 2027.
My research asks how single- and multi-agent reinforcement-learning methods can share reusable implementations without giving up experimental flexibility or training performance.
The work began with learned multi-agent communication, progressed through SARNet and a multiprocessing TensorFlow trainer, and led to casRL: a device-resident PyTorch framework for composing complete RL and MARL experiments.
2020–Present · Sole creator and principal architect
casRL is an open-source research framework that builds RL and MARL experiments from reusable implementations of policies, rollout collection, replay, temporal processing, objectives, optimization, communication, and model updates.
It implements more than 100 algorithms and variants across online, offline, hybrid, single-agent, and multi-agent learning. Researchers can configure synchronous and asynchronous rollout and training together with heterogeneous environments and datasets, safety constraints, planning, world models, auxiliary tasks, teacher–student learning, and federated learning.
The compiler validates a complete experiment before execution and the runtime prebinds its operations and data movement. casRL sustains 2–4 million environment steps per second during end-to-end single-node training on one CPU and one GPU.
Current research
I am developing field-topology feedback based on Gray-code Morton order for coordinating heterogeneous swarms over spatially structured fields.
The work studies how feedback between spatial structure and communication topology can support coordination as agents differ in sensing, capability, and role.
Murtaza Rangwala and R. K. Williams
Manuscript in preparation; targeting MLSys 2027.
Murtaza Rangwala and R. K. Williams
Manuscript in preparation; targeting JMLR.
Murtaza Rangwala and R. K. Williams
Manuscript in preparation; targeting ICRA 2027.
Jun Liu, Murtaza Rangwala, Kulbir S. Ahluwalia, Shayan Ghajar, Harnaik S. Dhami, Pratap Tokekar, Benjamin F. Tracy, Ryan K. Williams
IEEE T-ASE, 2024
A data-synthesis, prediction, planning, and intermittent-deployment pipeline for large-scale agricultural robot teams.
Murtaza Rangwala, Jun Liu, Kulbir S. Ahluwalia, Shayan Ghajar, Harnaik S. Dhami, Benjamin F. Tracy, Pratap Tokekar, Ryan K. Williams
Agronomy, 2021
Long-horizon pasture prediction from synthetic spatio-temporal datasets for planning agricultural-robot deployments.
Murtaza Rangwala, Ryan K. Williams
NeurIPS, 2020
A memory-based attention architecture that learns which inter-agent messages matter while reasoning over prior information.