Research

My research asks how single- and multi-agent reinforcement-learning methods can share reusable implementations without giving up experimental flexibility or training performance.

The work began with learned multi-agent communication, progressed through SARNet and a multiprocessing TensorFlow trainer, and led to casRL: a device-resident PyTorch framework for composing complete RL and MARL experiments.

2020–Present · Sole creator and principal architect

Coordination for Scale RL

casRL is an open-source research framework that builds RL and MARL experiments from reusable implementations of policies, rollout collection, replay, temporal processing, objectives, optimization, communication, and model updates.

It implements more than 100 algorithms and variants across online, offline, hybrid, single-agent, and multi-agent learning. Researchers can configure synchronous and asynchronous rollout and training together with heterogeneous environments and datasets, safety constraints, planning, world models, auxiliary tasks, teacher–student learning, and federated learning.

The compiler validates a complete experiment before execution and the runtime prebinds its operations and data movement. casRL sustains 2–4 million environment steps per second during end-to-end single-node training on one CPU and one GPU.

Current research

Field-topology feedback for heterogeneous swarms

I am developing field-topology feedback based on Gray-code Morton order for coordinating heterogeneous swarms over spatially structured fields.

The work studies how feedback between spatial structure and communication topology can support coordination as agents differ in sensing, capability, and role.

Manuscripts in preparation

casRL: A Compiler and Runtime for Composable Reinforcement Learning Systems

Murtaza Rangwala and R. K. Williams

Manuscript in preparation; targeting MLSys 2027.

casRL: Composable Systems for Unified Single- and Multi-Agent Reinforcement Learning

Murtaza Rangwala and R. K. Williams

Manuscript in preparation; targeting JMLR.

Field-Topology Feedback via Morton Z-Codes for Heterogeneous Swarm Coordination

Murtaza Rangwala and R. K. Williams

Manuscript in preparation; targeting ICRA 2027.

Publications

T-ASE

Intermittent Deployment for Large-Scale Multi-Robot Forage Perception: Data Synthesis, Prediction, and Planning

Jun Liu, Murtaza Rangwala, Kulbir S. Ahluwalia, Shayan Ghajar, Harnaik S. Dhami, Pratap Tokekar, Benjamin F. Tracy, Ryan K. Williams

IEEE T-ASE, 2024

A data-synthesis, prediction, planning, and intermittent-deployment pipeline for large-scale agricultural robot teams.

DP

DeepPaSTL: Spatio-Temporal Deep Learning Methods for Predicting Long-Term Pasture Terrains Using Synthetic Datasets

Murtaza Rangwala, Jun Liu, Kulbir S. Ahluwalia, Shayan Ghajar, Harnaik S. Dhami, Benjamin F. Tracy, Pratap Tokekar, Ryan K. Williams

Agronomy, 2021

Long-horizon pasture prediction from synthetic spatio-temporal datasets for planning agricultural-robot deployments.

SAR

Learning Multi-Agent Communication through Structured Attentive Reasoning

Murtaza Rangwala, Ryan K. Williams

NeurIPS, 2020

A memory-based attention architecture that learns which inter-agent messages matter while reasoning over prior information.