ADAPTING MULTIMODAL SYSTEMS TO INFERENCE TIME VARIATIONS �
Jason Wu
Final Oral Defense for Ph.D. in Electrical and Computer Engineering
Advisor and Committee Chair: Prof. Mani Srivastava
Committee Members: Prof. Suhas Diggavi, Prof. Cho-Jui Hsieh, Prof. Jonathan Kao
1
Multimodal Systems
2
Autonomous Driving
Disaster Response
Fuse information from two or more distinct modalities to accomplish an objective
What is a Modality?
Modality: A unique mapping from the physical world to a sensor space via a transducer
3
Distributed Multimodal System
Camera
Radar
Camera
Multimodal Deep Neural Networks
Intra-Modality and Inter-Modality Features
4
Fusion
Why Multimodal?
Enhanced Performance:
5
Autonomous Driving
Disaster Response
Increased Robustness
Inference-Time Variations
6
Multimodal systems suffer from performance degradation from inference time variations
Misconception of Multimodal Systems
7
[1] D. Hazarika, Y. Li, B. Cheng, S. Zhao, R. Zimmermann, and S. Poria, “Analyzing modality robustness in multimodal sentiment analysis,” arXiv preprint arXiv:2205.15465, 2022.
How do we minimize the performance degradation from inference time variations to make multimodal networks suitable for real-world environments?
General Framework
8
Adapt the weights or architecture of the multimodal neural network according to information collected about the world to mitigate inference-time variations
Key Idea
Models and Tasks
9
3D Detection
VQA
Single-Instance, Task-Specific Neural Networks
Multi-Instance, Task-Specific Neural Networks
Foundation Models
Gesture Detection
Thesis Outline
10
Task Specific Models
Foundation Models
Localization
Deployment Time
Sensor Perspective Shift
Localization
Classification
Runtime
Varying Modality Quality; Compute Resources
3D Detection in AVs
Runtime
Varying Modality Quality; Compute Resources; Sample Complexity
VQA
Deployment Time
Highly Specialized Domains
Chapter 2
11
Chapter 3
12
Chapter 4
13
Chapter 5
14
FlexLoc: Conditional Neural Networks for Zero-Shot Sensor Perspective Invariance in Object Localization with Distributed Multimodal Sensors
15
IROS 2024
Pre-Quals
Wu, J., Wang, Z., Ouyang, X., Jeong, H. L., Samplawski, C., Kaplan, L. M., ... & Srivastava, M. (2024, October). FlexLoc: Conditional Neural Networks for Zero-Shot Sensor Perspective Invariance in Object Localization with Distributed Multimodal Sensors. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (pp. 8563-8570). IEEE.
Motivation
Infrastructure-based localization is a critical technology
16
Smart City Applications
Delivery Robots
Multimodal Localization
17
Camera
Camera
Radar
Unimodal System
Multimodal System
Distributed Multimodal Localization
18
Multimodal Node
Multimodal Node
Multimodal Node
Single-View System
Multi-View System
Deployment Time Sensor Perspective Shift
19
Motivational Experiment
Localizing a small car in a 3m x 3m track
20
Related Work -- Datasets
21
C. Samplawski, S. Fang, Z. Wang, D. Ganesan, M. Srivastava, and B. M. Marlin, “Heteroskedastic geospatial tracking with distributed camera networks,” in Proceedings of the Thirty-Ninth Conference on Uncertainty in Artificial Intelligence, ser. UAI ’23. JMLR.org, 2023.
M. Salimibeni, Z. Hajiakhondi-Meybodi, P. Malekzadeh, M. Atashi, K. N. Plataniotis, and A. Mohammadi, “Iot-td: Iot dataset for multiple model ble-based indoor localization/tracking,” in 2020 28th European Signal Processing Conference (EUSIPCO), 2021, pp. 1697–1701.
Z. Kandylakis, K. Vasili, and K. Karantzalos, “Fusing multimodal video data for detecting moving objects/targets in challenging indoor and outdoor scenes,” Remote Sensing, vol. 11, no. 4, p. 446, 2019.
K. Nakamura, K. Nakadai, F. Asano, and G. Ince, “Intelligent sound source localization and its application to multimodal human tracking,” in 2011 IEEE/RSJ International Conference on Intelligent Robots and Systems, 2011, pp. 143–148
H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom, “nuscenes: A multimodal dataset for autonomous driving,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 11 621–11 631.
P. Sun, H. Kretzschmar, X. Dotiwalla, A. Chouard, V. Patnaik, P. Tsui, J. Guo, Y. Zhou, Y. Chai, B. Caine, et al., “Scalability in perception for autonomous driving: Waymo open dataset,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 2446–2454.
M. J. Bocus, W. Li, S. Vishwakarma, R. Kou, C. Tang, K. Woodbridge, I. Craddock, K. McConville, et al., “Operanet, a multimodal activity recognition dataset acquired from radio frequency and vision-based sensors,” Scientific data, vol. 9, no. 1, p. 474, 2022.
C. Torres, J. C. Fried, K. Rose, and B. S. Manjunath, “A multiview multimodal system for monitoring patient sleep,” IEEE Transactions on Multimedia, vol. 20, no. 11, pp. 3057–3068, 2018.
Salas-Moreno, R. F., Newcombe, R. A., Strasdat, H., Kelly, P. H., & Davison, A. J. (2013). Slam++: Simultaneous localisation and mapping at the level of objects. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 1352-1359).
Multimodal, Multiview localization systems are underexplored
Related Work
22
M. J. Bocus, W. Li, S. Vishwakarma, R. Kou, C. Tang, K. Woodbridge, I. Craddock, K. McConville, et al., “Operanet, a multimodal activity recognition dataset acquired from radio frequency and vision-based sensors,” Scientific data, vol. 9, no. 1, p. 474, 2022.
C. Torres, J. C. Fried, K. Rose, and B. S. Manjunath, “A multiview multimodal system for monitoring patient sleep,” IEEE Transactions on Multimedia, vol. 20, no. 11, pp. 3057–3068, 2018.
Salas-Moreno, R. F., Newcombe, R. A., Strasdat, H., Kelly, P. H., & Davison, A. J. (2013). Slam++: Simultaneous localisation and mapping at the level of objects. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 1352-1359).
S. Yang, L. Hou, H. Huang, C. Ma, P. Wan, D. Zhang, X. Chen, and J. Liao, “Direct-a-video: Customized video generation with user-directed camera movement and object motion,” arXiv preprint arXiv:2402.03162, 2024.
Y. Zhao, S. Kong, and C. Fowlkes, “Camera pose matters: Improving depth prediction by mitigating pose distribution bias,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 15 759–15 768.
Existing work does not address sensor perspective shift in multimodal, multiview localization systems
Technique | Papers | Drawbacks |
Multimodal, Multiview Infrastructure-Based Localization | Bocus et al. Torres et al. | Limited modalities No perspective shift |
Perspective Invariance in SLAM | Salas Moreno et al. | Egocentric Only Limited Modalities Reliant on Scan Matching |
Pose Injection in Other Setting | Yang et al. Zhao et al. | Not related to localization |
Proposed Solution
23
Inject knowledge of the world state (i.e., sensor poses) into a localization neural network
FlexLoc Main Contributions
24
Conditional Neural Networks
25
Conditional Neural Networks
26
During training, the controller learns how sensor pose impact the localization result
Base Architecture
27
Conditional Convolution (CondConv)
28
Conditional Layer/Batch Normalization
29
Results
30
Dataset
31
Jeong, H. L., Wang, Z., Samplawski, C., Wu, J., Fang, S., Kaplan, L. M., ... & Srivastava, M. (2024). Gdtm: An indoor geospatial tracking dataset with distributed multimodal sensors. arXiv preprint arXiv:2402.14136.
Baselines
32
FlexLoc vs. Baselines
33
FlexLoc outperforms baselines by almost 50%
Euclidean Distance Error (cm) on GDTM
FlexLoc vs Baselines
34
Low 90th percentile error
FlexLoc vs Baselines
35
Ablation Study: Pose Injection
36
FlexLoc Overhead
37
Extended Evaluations
38
Discussions and Future Work
39
Transition to Runtime Variations
40
ADMN: A Layer-Wise Adaptive Multimodal Network for Dynamic Input Noise and Compute Resources
41
Neurips 2025
Pre-Quals
Wu, J., Yuan, Y., Yang, K., Kaplan, L., & Srivastava, M. (2025). ADMN: A Layer-Wise Adaptive Multimodal Network for Dynamic Input Noise and Compute Resources. Advances in Neural Information Processing Systems
Motivation - Varying Modality Quality
42
Dynamic Modality Quality of Information (QoI)
Motivation – Computational Resources
43
Runtime
Challenges
44
Related Work – Early-Exit
45
J. Xin, R. Tang, J. Lee, Y. Yu, and J. Lin, “DeeBERT: Dynamic early exiting for accelerating BERT inference,” arXiv preprint arXiv:2004.12993, 2020.
W. Zhou, C. Xu, T. Ge, J. McAuley, K. Xu, and F. Wei, “BERT Loses Patience: Fast and Robust Inference with Early Exit,” in Advances in Neural Information Processing Systems, vol. 33, 2020, pp. 18 330–18 341.
Drawbacks:
Related Work – Dynamic Multimodal Inference
46
Drawbacks:
Xue, Z., & Marculescu, R. (2023). Dynamic multimodal fusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 2575-2584).
Alikhani, H., Kanduri, A., Liljeberg, P., Rahmani, A. M., & Dutt, N. (2023). DynaFuse: dynamic fusion for resource efficient multimodal machine learning inference. IEEE Embedded Systems Letters, 15(4), 222-225.
Panda, R., Chen, C. F. R., Fan, Q., Sun, X., Saenko, K., Oliva, A., & Feris, R. (2021). Adamml: Adaptive multi-modal learning for efficient video recognition. In Proceedings of the IEEE/CVF international conference on computer vision (pp. 7576-7585).
Gao, R., Oh, T. H., Grauman, K., & Torresani, L. (2020). Listen to look: Action recognition by previewing audio. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 10457-10467)
Cai, Q., Liu, X., Zhang, K., Xie, X., Tong, X., & Li, K. (2023). Acf: An adaptive compression framework for multimodal network in embedded devices. IEEE Transactions on Mobile Computing, 23(5), 5195-5211.
.
Related Work – Unimodal Subnetworks
47
Cai, H., Gan, C., Wang, T., Zhang, Z., & Han, S. (2019). Once-for-all: Train one network and specialize it for efficient deployment. arXiv preprint arXiv:1908.09791.
Fan, A., Grave, E., & Joulin, A. (2019). Reducing transformer depth on demand with structured dropout. arXiv preprint arXiv:1909.11556.
Proposed Solution
48
ADMN Main Contributions
49
ADMN Overall Architecture
50
ADMN Overall Architecture
51
Controller Architecture
52
QoI Supervision
53
Results
54
Datasets and Corruptions
55
Dataset | Attributes |
GDTM | RGB, Depth Localization |
MM-Fi | RGB, Depth HAR |
AVE | Audio, RGB Event videos |
Corruptions | Modalities/Datasets |
Gaussian Noise | RGB, Depth in GDTM, MMFI |
Lowlight | RGB in GDTM, AVE |
Blur | RGB in GDTM |
Background Audio | Audio in AVE |
Rain | RGB in AVE |
Baselines
56
Method | Explanation |
Upper Bound | Allocate all layers (i.e., 12) to each backbone |
Naïve Allocation | |
X Modality Only | |
Naïve Scratch | |
Modality Network Selection (MNS) | Train several expert models for every budget, and use a controller to select among them |
Allocation Baselines: Upper Bound, Naïve Allocation, X Modality Only
New Training Baselines: Naïve Scratch, MNS
Main Results – GDTM Localization
57
Upper Bound utilizes 24 layers (12 for each modality backbone)
Main Results – Classification
58
ADMN Overhead
59
GDTM:
MM-Fi:
Discussions and Future Work
60
SWAN: World-Aware Adaptive Multimodal Networks for Runtime Variations
61
ECCV 2026
Wu, J., Jin, S.S., Yuan, Y., Wigness, M., Kaplan, L.M., Qiu, H., Srivastava, M.: SWAN: World-Aware Adaptive Multimodal Networks for Runtime Variations. In: Proceedings of the 19th European Conference on Computer Vision (ECCV). (2026)
Motivation
62
Related Work
63
Existing Multimodal AV Deep Learning Networks neglect to consider adaptation to runtime variations
Example from nuScenes
Variable Sample Complexity
64
Challenges
65
SWAN Contributions
66
SWAN Design
67
SWAN Architecture
68
SWAN Controller
69
NeuralSort Example
70
NeuralSort
Exploration
Exploitation
SWAN SkipGate Module
71
Jointly tackling several variations requires careful module design
SWAN Token Pruning
72
Results
73
Datasets and Baselines
74
MultiCorrupt darkness
Main Results
75
NDS and mAP values on Multicorrupt nuScenes, C: Controller, S: SkipGate, P: Pruning
Evaluated on RTX 4090
Layer Allocations
76
Layer allocations across modality for controller and SkipGate on MultiCorrupt nuScenes
Practical Considerations
77
Jetson Orin AGX
78
SWAN’s savings are more pronounced on weaker hardware
TensorRT Deployment
79
SWAN is more effective on edge hardware and with TensorRT runtime!
Localization Error (cm) on the GDTM dataset
Inferring Available Compute
80
Proof-of-Concept System (Redraw Fig)
81
TensorRT model on Jetson Orin AGX; Background GPU process alternating HIGH (1900) and LOW (2); 30 ms latency threshold
Conclusion and Discussion
82
Decoupling Vision and Language: Codebook Anchored Visual Adaptation
83
CVPR 2026
J. Wu, T. Zhao, C. Liu, J. Cai, Z. Zhang, Z. Li, A. Singh, X. Xu, M. Srivastava, and J. Wu, "Decoupling Vision and Language: Codebook Anchored Visual Adaptation," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026.
Background
84
Background - LVLMs
85
LVLMs “glue” a pretrained vision encoder and LLM together
Background – Discrete LVLMs
86
Vision
Encoder
Codebook
LLM
Tabby
Cat
Tell me the species of this cat and give it a tie
Forward pass of Discrete LVLM
Challenges
87
VILA-U 7B output on Plant Disease Identification (PlantVillage)
Existing Solutions
88
[1] Gregor Geigle, Radu Timofte, and Goran Glavaš. African or european swallow? benchmarking large visionlanguage models for fine-grained object classification. arXiv:2406.14496, 2024
[2] Jiawei Chen, Dingkang Yang, Yue Jiang, Mingcheng Li, Jinjie Wei, Xiaolu Hou, and Lihua Zhang. Efficiency in focus: Layernorm as a catalyst for fine-tuning medical visual language models. In ACM International Conference on Multimedia, 2024.
[3] Llava-radz: Can multimodal large language models effectively tackle zero-shot radiology recognition?
Can we finetune the vision encoder without involving the massive LLM?
Proposed Solution
89
Intuition
90
Vision
Encoder
Codebook
LLM
1B
Bengal
Cat
Finetuned
Vision
Encoder
Codebook
LLM
1B
Tabby
Cat
Before Fine-Tuning
After Fine-Tuning
Intuition
91
Substitute Vision Encoder Finetuned with 1B LLM
CRAFT Key Contributions
92
CRAFT
93
Losses
94
CRAFT: Token Pruning
95
Results
96
Datasets and Baselines
97
Main Results
Surrogate Fine-tuning vs. Baselines averaged across 10 benchmarks
98
Quality of Generated Output
99
Correctness, Presence, Relevance, Faithfulness, Verbosity of Explanations
Cross-Compatibility
100
Cross-LLM Transfer, accuracy averaged across 10 benchmarks
Training Efficiency
101
Training on Stanford-Dogs Dataset for 1 Epoch
Discussion and Conclusions
102
THESIS CONCLUSION
103
Thesis Conclusion
104
Future Work (1)
Runtime Adaptation of Foundation Models
105
Future Work (2)
Physical Sensor Adaptation
106
Acknowledgements
Advisor: Mani Srivastava
Funding Agencies: CONIX, IoBT, NDSEG
NESL Collaborators: Ziqi Wang, Xiaomin Ouyang, Yuyang Yuan, Ankur Sarker, Lucas Jeong
External Collaborators: Lance Kaplan (ARL), Benjamin Marlin (UMass), Colin Samplawski (UMass)
107
THANK YOU
108
Backup – Alternative Multimodal Structures
109
Backup – Conditional Batch Normalization
110
Modulating early visual processing by language
Neurips 2017
Backup – Gumbel Softmax
111
Backup – ADMN Corruptions
112
Backup – ADMN Three Modalities
113
Backup – ADMN Train from Scratch
114
Backup – ADMN Visual Results
115
Backup – ADMN Visual Results
116
Backup – ADMN Universal Controller
117
SWAN vs ADMN
118
SWAN Visual Example
119
SWAN Pruning
120
SWAN Latency Based
121
SWAN More Datasets
122
CRAFT Full Table 1
123
CRAFT Language
124
CRAFT Loss Ablation
125
CRAFT Explanations
126
CRAFT Cross Compat
127
CRAFT Discrete Indices
128
Figures/Content
129
Wu, J., Yuan, Y., Yang, K., Kaplan, L., & Srivastava, M. (2025). ADMN: A Layer-Wise Adaptive Multimodal Network for Dynamic Input Noise and Compute Resources. Advances in Neural Information Processing Systems
130
Multimodal Network
Multimodal Inputs
World Info
Sensor Pose
Modality Quality
Platform Status
New Data
Adapt the weights or architecture of the multimodal neural network according to information collected about the world to mitigate inference-time variations
Key Idea
131
Task Specific Models
Foundation Models
Localization
Deployment Time
Sensor Perspective Shift
Localization
Classification
Runtime
Varying Modality Quality; Compute Resources
3D Detection in AVs
Runtime
Varying Modality Quality; Compute Resources; Sample Complexity
VQA
Deployment Time
Highly Specialized Domains
132
Task Specific Models
Foundation Models
Localization
Deployment Time
Sensor Perspective Shift
Localization
Classification
Runtime
Varying Modality Quality; Compute Resources
3D Detection in AVs
Runtime
Varying Modality Quality; Compute Resources; Sample Complexity
VQA
Deployment Time
Highly Specialized Domains
133
Task Specific Models
Foundation Models
Localization
Deployment Time
Sensor Perspective Shift
Localization
Classification
Runtime
Varying Modality Quality; Compute Resources
3D Detection in AVs
Runtime
Varying Modality Quality; Compute Resources; Sample Complexity
VQA
Deployment Time
Highly Specialized Domains
Problem Space
134
Task Specific Models
Foundation Models
Localization
Deployment Time
Sensor Perspective Shift
Localization
Classification
Runtime
Varying Modality Quality; Compute Resources
3D Detection in AVs
Runtime
Varying Modality Quality; Compute Resources; Sample Complexity
VQA
Deployment Time
Highly Specialized Domains
135