Sitemap

A list of all the posts and pages found on the site. For you robots out there, there is an XML version available for digesting as well.

Pages

Posts

Future Blog Post

less than 1 minute read

Published:

This post will show up by default. To disable scheduling of future posts, edit config.yml and set future: false.

Blog Post number 4

less than 1 minute read

Published:

This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.

Blog Post number 3

less than 1 minute read

Published:

This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.

Blog Post number 2

less than 1 minute read

Published:

This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.

Blog Post number 1

less than 1 minute read

Published:

This is a sample blog post. Lorem ipsum I can’t remember the rest of lorem ipsum and don’t have an internet connection right now. Testing testing testing this blog post. Blog posts are cool.

portfolio

publications

Interact: Advancing large-scale versatile 3d human-object interaction generation

Published in CVPR 2025, 2025

InterAct introduces a large-scale benchmark for 3D human-object interaction generation, addressing the limitations of existing HOI datasets such as limited scale, inconsistent annotations, and artifacts like penetration or floating contacts. The work consolidates and standardizes diverse HOI data, enriches it with detailed textual annotations, and proposes an optimization-based framework for interaction correction and augmentation. It further defines six HOI generation tasks and demonstrates that unified contact-aware modeling improves performance across text-conditioned, action-conditioned, prediction, and imitation settings.

Recommended citation: Sirui Xu, Dongting Li, Yucheng Zhang, Xiyan Xu, Qi Long, Ziyin Wang, Yunzhi Lu, Shuchang Dong, Hezi Jiang, Akshat Gupta, Yu-Xiong Wang, Liang-Yan Gui
Download Paper

Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective

Published in ICML 2026, 2026

This paper studies inference-time alignment under KL regularization, where directly optimizing a reward model learned from user preferences can be suboptimal because the base policy may bias the aligned model away from the user’s intended utility. We formulate reward model design as a Stackelberg game between the reward provider and the LLM policy, and derive a threshold-based reward shaping method that approximates the optimal reward model. Our approach integrates into existing inference-time alignment methods with minimal overhead and consistently improves reward while maintaining comparable diversity and coherence.

Recommended citation: Wang, H., Lin, T., Kong, L., Li, C., Jiang, H., & Tambe, M. Reward Shaping for (Inference-Time) Alignment: A Stackelberg Game Perspective. ICML 2026.
Download Paper

Latent Spherical Flow Policy for Reinforcement Learning with Combinatorial Actions

Published in ICML 2026 Spotlight, 2026

This paper studies reinforcement learning with combinatorial action spaces, where feasible actions are exponentially large and constrained by hard structural requirements. We propose LSFLOW, a latent spherical flow policy that learns an expressive stochastic distribution over continuous cost directions and uses a combinatorial optimization solver to map each sample to a valid structured action. By training the critic directly in the latent space and introducing a smoothed Bellman operator to handle solver-induced discontinuities, our method improves both performance and efficiency across combinatorial RL benchmarks and a real-world STI testing application.

Recommended citation: Kong, L., Satish, A., Jiang, H., Kangaslahti, A., Ma, A., Chen, W., Song, M., Xu, L., & Tambe, M. Latent Spherical Flow Policy for Reinforcement Learning with Combinatorial Actions. ICML 2026 Spotlight.
Download Paper

Generative Frontier Planning for Adaptive Peer-Referral Recruitment under Covariate-Dependent Arrivals

Published in KDD 2026 Workshop on epiDAMIK, 2026

This paper studies adaptive peer-referral recruitment, where limited referral resources must be allocated across multiple rounds and current decisions affect both the number and covariates of future recruits. We propose Generative Frontier Planning, a model-based planner that learns covariate-dependent referral dynamics and uses a latent coverage value surrogate to make future-frontier planning tractable. By replacing per-step Monte Carlo sampling with deterministic backups and exploiting diminishing returns for greedy allocation, our method improves recruitment performance over random, reinforcement-learning, and population-level dynamic-programming baselines.

Recommended citation: Kong, L.*, Jiang, H.*, Ma, A., Wang, K., Kangaslahti, A., & Tambe, M. Generative Frontier Planning for Adaptive Peer-Referral Recruitment under Covariate-Dependent Arrivals. KDD 2026 Workshop epiDAMIK. (*Equal contribution)
Download Paper

talks

teaching

Teaching experience 1

Undergraduate course, University 1, Department, 2014

This is a description of a teaching experience. You can use markdown like any other post.

Teaching experience 2

Workshop, University 1, Department, 2015

This is a description of a teaching experience. You can use markdown like any other post.