1 of 135

ADAPTING MULTIMODAL SYSTEMS TO INFERENCE TIME VARIATIONS �

Jason Wu

Final Oral Defense for Ph.D. in Electrical and Computer Engineering

Advisor and Committee Chair: Prof. Mani Srivastava

Committee Members: Prof. Suhas Diggavi, Prof. Cho-Jui Hsieh, Prof. Jonathan Kao

1

2 of 135

Multimodal Systems

2

Autonomous Driving

Disaster Response

Fuse information from two or more distinct modalities to accomplish an objective

3 of 135

What is a Modality?

Modality: A unique mapping from the physical world to a sensor space via a transducer

3

Distributed Multimodal System

Camera

Radar

Camera

4 of 135

Multimodal Deep Neural Networks

Intra-Modality and Inter-Modality Features

4

 

 

 

 

Fusion

 

5 of 135

Why Multimodal?

Enhanced Performance:

    • Intra and Inter-Modality features result in improved performance

5

Autonomous Driving

Disaster Response

Increased Robustness

    • Corruption of one modality’s data does not irrecoverably damage system’s accuracy

6 of 135

Inference-Time Variations

  • Highly dynamic environments where multimodality is relevant contain inference-time variations
  • Deployment Time:
    • New sensor locations, environmental mediums, computing platforms, or tasks
  • Runtime:
    • Shifting modality quality, unstable computational resources

6

Multimodal systems suffer from performance degradation from inference time variations

7 of 135

Misconception of Multimodal Systems

  • How can multimodal networks simultaneously be more robust yet brittle?
  • Existing literature focuses largely on a narrow slice of robustness: with respect to sensor noise and for metric of accuracy [1]
  • Inference-time variations go beyond sensor noise and impact system performance beyond accuracy (e.g., latency)
  • Increased number of modalities causes greater vulnerability
    • Deployment time sensor shifts
    • Runtime resource allocation

7

[1] D. Hazarika, Y. Li, B. Cheng, S. Zhao, R. Zimmermann, and S. Poria, “Analyzing modality robustness in multimodal sentiment analysis,” arXiv preprint arXiv:2205.15465, 2022.

How do we minimize the performance degradation from inference time variations to make multimodal networks suitable for real-world environments?

8 of 135

General Framework

8

Adapt the weights or architecture of the multimodal neural network according to information collected about the world to mitigate inference-time variations

Key Idea

9 of 135

Models and Tasks

9

3D Detection

VQA

Single-Instance, Task-Specific Neural Networks

  • Only a single sensing target in each sample
  • Trained for one specific task

Multi-Instance, Task-Specific Neural Networks

  • Multiple sensing targets in each sample
  • Trained for one specific task

Foundation Models

  • Trained on internet scale data (single and multi-instance tasks)
  • Tested on arbitrary tasks

Gesture Detection

10 of 135

Thesis Outline

10

Task Specific Models

Foundation Models

Localization

Deployment Time

Sensor Perspective Shift

Localization

Classification

Runtime

Varying Modality Quality; Compute Resources

3D Detection in AVs

Runtime

Varying Modality Quality; Compute Resources; Sample Complexity

VQA

Deployment Time

Highly Specialized Domains

11 of 135

Chapter 2

11

  • Sensor nodes can change position and orientation during deployment
  • Conditioning the neural network on sensor pose provides robustness

12 of 135

Chapter 3

12

  • Runtime variations occur on a faster time-scale
    • Fluctuating modality quality
    • Dynamic computational resource availability
  • Neural Network modulates total resource usage and allocation across modalities

13 of 135

Chapter 4

13

  • Autonomous Driving: Complex multi-object 3D detection
  • New Considerations:
    • Fluctuating modality quality
    • Dynamic computational resource availability
    • Variable sample complexity (e.g., more or less targets)

14 of 135

Chapter 5

14

  • Adapting Large Vision-Language models to highly specialized domains
  • Large size of these models and asymmetric parameter count across modalities creates unique challenges
  • Perform memory efficient visual fine-tuning for easy adaptation

15 of 135

FlexLoc: Conditional Neural Networks for Zero-Shot Sensor Perspective Invariance in Object Localization with Distributed Multimodal Sensors

15

IROS 2024

Pre-Quals

Wu, J., Wang, Z., Ouyang, X., Jeong, H. L., Samplawski, C., Kaplan, L. M., ... & Srivastava, M. (2024, October). FlexLoc: Conditional Neural Networks for Zero-Shot Sensor Perspective Invariance in Object Localization with Distributed Multimodal Sensors. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (pp. 8563-8570). IEEE.

16 of 135

Motivation

Infrastructure-based localization is a critical technology

16

Smart City Applications

Delivery Robots

17 of 135

Multimodal Localization

17

Camera

Camera

Radar

Unimodal System

Multimodal System

18 of 135

Distributed Multimodal Localization

18

Multimodal Node

Multimodal Node

Multimodal Node

Single-View System

Multi-View System

19 of 135

Deployment Time Sensor Perspective Shift

19

20 of 135

Motivational Experiment

Localizing a small car in a 3m x 3m track

20

21 of 135

Related Work -- Datasets

  • Unimodal, Multi-View Localization Infrastructure: Samplawski et al., Salimibeni et al.
  • Multimodal, Single View Localization Infrastructure: Kandylakis et al., Nakamura et al.
  • Egocentric Localization: Caesar et al., Sun et al.
  • Multimodal, multi-view Infrastructure (Bocus et al., Torres et al.): Limited modalities, no perspective shift

21

C. Samplawski, S. Fang, Z. Wang, D. Ganesan, M. Srivastava, and B. M. Marlin, “Heteroskedastic geospatial tracking with distributed camera networks,” in Proceedings of the Thirty-Ninth Conference on Uncertainty in Artificial Intelligence, ser. UAI ’23. JMLR.org, 2023.

M. Salimibeni, Z. Hajiakhondi-Meybodi, P. Malekzadeh, M. Atashi, K. N. Plataniotis, and A. Mohammadi, “Iot-td: Iot dataset for multiple model ble-based indoor localization/tracking,” in 2020 28th European Signal Processing Conference (EUSIPCO), 2021, pp. 1697–1701.

Z. Kandylakis, K. Vasili, and K. Karantzalos, “Fusing multimodal video data for detecting moving objects/targets in challenging indoor and outdoor scenes,” Remote Sensing, vol. 11, no. 4, p. 446, 2019.

K. Nakamura, K. Nakadai, F. Asano, and G. Ince, “Intelligent sound source localization and its application to multimodal human tracking,” in 2011 IEEE/RSJ International Conference on Intelligent Robots and Systems, 2011, pp. 143–148

H. Caesar, V. Bankiti, A. H. Lang, S. Vora, V. E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom, “nuscenes: A multimodal dataset for autonomous driving,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 11 621–11 631.

P. Sun, H. Kretzschmar, X. Dotiwalla, A. Chouard, V. Patnaik, P. Tsui, J. Guo, Y. Zhou, Y. Chai, B. Caine, et al., “Scalability in perception for autonomous driving: Waymo open dataset,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 2446–2454.

M. J. Bocus, W. Li, S. Vishwakarma, R. Kou, C. Tang, K. Woodbridge, I. Craddock, K. McConville, et al., “Operanet, a multimodal activity recognition dataset acquired from radio frequency and vision-based sensors,” Scientific data, vol. 9, no. 1, p. 474, 2022.

C. Torres, J. C. Fried, K. Rose, and B. S. Manjunath, “A multiview multimodal system for monitoring patient sleep,” IEEE Transactions on Multimedia, vol. 20, no. 11, pp. 3057–3068, 2018.

Salas-Moreno, R. F., Newcombe, R. A., Strasdat, H., Kelly, P. H., & Davison, A. J. (2013). Slam++: Simultaneous localisation and mapping at the level of objects. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 1352-1359).

Multimodal, Multiview localization systems are underexplored

22 of 135

Related Work

22

M. J. Bocus, W. Li, S. Vishwakarma, R. Kou, C. Tang, K. Woodbridge, I. Craddock, K. McConville, et al., “Operanet, a multimodal activity recognition dataset acquired from radio frequency and vision-based sensors,” Scientific data, vol. 9, no. 1, p. 474, 2022.

C. Torres, J. C. Fried, K. Rose, and B. S. Manjunath, “A multiview multimodal system for monitoring patient sleep,” IEEE Transactions on Multimedia, vol. 20, no. 11, pp. 3057–3068, 2018.

Salas-Moreno, R. F., Newcombe, R. A., Strasdat, H., Kelly, P. H., & Davison, A. J. (2013). Slam++: Simultaneous localisation and mapping at the level of objects. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 1352-1359).

S. Yang, L. Hou, H. Huang, C. Ma, P. Wan, D. Zhang, X. Chen, and J. Liao, “Direct-a-video: Customized video generation with user-directed camera movement and object motion,” arXiv preprint arXiv:2402.03162, 2024.

Y. Zhao, S. Kong, and C. Fowlkes, “Camera pose matters: Improving depth prediction by mitigating pose distribution bias,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 15 759–15 768.

Existing work does not address sensor perspective shift in multimodal, multiview localization systems

Technique

Papers

Drawbacks

Multimodal, Multiview Infrastructure-Based Localization

Bocus et al.

Torres et al.

Limited modalities

No perspective shift

Perspective Invariance in SLAM

Salas Moreno et al.

Egocentric Only

Limited Modalities

Reliant on Scan Matching

Pose Injection in Other Setting

Yang et al.

Zhao et al.

Not related to localization

23 of 135

Proposed Solution

23

Inject knowledge of the world state (i.e., sensor poses) into a localization neural network

24 of 135

FlexLoc Main Contributions

  • First exploration into applying conditional neural networks to address test-time robustness in multimodal, multi-view localization
  • Minimal additional parameters (< 0.2% of model parameters), with a simple implementation allowing for easy integration

24

25 of 135

Conditional Neural Networks

25

26 of 135

Conditional Neural Networks

26

During training, the controller learns how sensor pose impact the localization result

27 of 135

Base Architecture

27

28 of 135

Conditional Convolution (CondConv)

28

29 of 135

Conditional Layer/Batch Normalization

29

30 of 135

Results

30

31 of 135

Dataset

  • GDTM Dataset: Nine hours of data with 22 unique viewpoints
    • Training: 13 views (105 min)
    • Validation: 4 views (35 min)
    • Testing: 5 views (55 min)

31

Jeong, H. L., Wang, Z., Samplawski, C., Wu, J., Fang, S., Kaplan, L. M., ... & Srivastava, M. (2024). Gdtm: An indoor geospatial tracking dataset with distributed multimodal sensors. arXiv preprint arXiv:2402.14136.

32 of 135

Baselines

  • Unconditional Early Fusion: Early fusion with no conditional layers
  • Late Fusion Local Coordinate Transform: Each modality in each node predicts in local coordinate space

32

33 of 135

FlexLoc vs. Baselines

33

FlexLoc outperforms baselines by almost 50%

Euclidean Distance Error (cm) on GDTM

34 of 135

FlexLoc vs Baselines

34

Low 90th percentile error

35 of 135

FlexLoc vs Baselines

35

36 of 135

Ablation Study: Pose Injection

  • Switch architecture to ViT-style transformer backbones
  • Baselines: Cross-Attention (Yang et al.), Input Concatenation (Zhao et al.)

36

37 of 135

FlexLoc Overhead

  • Number of Parameters
  • Multiply and Accumulate Operations (MACs)

37

38 of 135

Extended Evaluations

  • Outdoor Dataset: GQ-Husky collected with AprilTag provided node poses
    • 100 square meters vs 9 square meters (GDTM)
    • 50 minutes of data with 26 perspectives (22 train, 4 test)

38

39 of 135

Discussions and Future Work

  • Summary: FlexLoc effectively mitigates sensor perspective shift through injection of node pose information
  • Extension to runtime variations:
    • Distributed sensors can be mobile during run time

39

40 of 135

Transition to Runtime Variations

  • Summary: FlexLoc mitigates sensor perspective shift with conditional neural networks
  • Extend the theme of model adaptation to runtime variations
  • Encompass a larger set of datasets across different tasks (e.g., classification, localization)

40

41 of 135

ADMN: A Layer-Wise Adaptive Multimodal Network for Dynamic Input Noise and Compute Resources

41

Neurips 2025

Pre-Quals

Wu, J., Yuan, Y., Yang, K., Kaplan, L., & Srivastava, M. (2025). ADMN: A Layer-Wise Adaptive Multimodal Network for Dynamic Input Noise and Compute Resources. Advances in Neural Information Processing Systems

42 of 135

Motivation - Varying Modality Quality

42

Dynamic Modality Quality of Information (QoI)

43 of 135

Motivation – Computational Resources

43

Runtime

44 of 135

Challenges

  • Existing multimodal networks are largely static
  • Modalities are allocated the same resources regardless of world state
  • Cannot accommodate variable computational resource availability

44

45 of 135

Related Work – Early-Exit

  • Early-Exit techniques adapt network behavior according to the input (easier samples exit)
  • DeeBert, PABEE: Confidence threshold to halt computation for simpler inputs

45

J. Xin, R. Tang, J. Lee, Y. Yu, and J. Lin, “DeeBERT: Dynamic early exiting for accelerating BERT inference,” arXiv preprint arXiv:2004.12993, 2020.

W. Zhou, C. Xu, T. Ge, J. McAuley, K. Xu, and F. Wei, “BERT Loses Patience: Fast and Robust Inference with Early Exit,” in Advances in Neural Information Processing Systems, vol. 33, 2020, pp. 18 330–18 341.

Drawbacks:

  • Neglect to consider dynamic QoI even in the unimodal setting, much less complex interplay among multimodal samples

46 of 135

Related Work – Dynamic Multimodal Inference

  • DynMM and DynaFuse: Train several expert networks with different modalities and perform network selection
  • AdaMML and Listen To Look: Use multimodal information to eliminate temporal redundancy in videos
  • ACF: Dynamically replace certain modules with lightweight networks

46

Drawbacks:

  • Overlook significant QoI variations arising from modality dependent corruption
  • Lack of fine-grained control over modality usage
  • Fail to consider varying resource budgets

Xue, Z., & Marculescu, R. (2023). Dynamic multimodal fusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 2575-2584).

Alikhani, H., Kanduri, A., Liljeberg, P., Rahmani, A. M., & Dutt, N. (2023). DynaFuse: dynamic fusion for resource efficient multimodal machine learning inference. IEEE Embedded Systems Letters15(4), 222-225.

Panda, R., Chen, C. F. R., Fan, Q., Sun, X., Saenko, K., Oliva, A., & Feris, R. (2021). Adamml: Adaptive multi-modal learning for efficient video recognition. In Proceedings of the IEEE/CVF international conference on computer vision (pp. 7576-7585).

Gao, R., Oh, T. H., Grauman, K., & Torresani, L. (2020). Listen to look: Action recognition by previewing audio. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 10457-10467)

Cai, Q., Liu, X., Zhang, K., Xie, X., Tong, X., & Li, K. (2023). Acf: An adaptive compression framework for multimodal network in embedded devices. IEEE Transactions on Mobile Computing23(5), 5195-5211.

.

47 of 135

Related Work – Unimodal Subnetworks

  • Address deployment-time compute variations by adjusting a single, large network
  • Once-For-All: Prune a large network across depth, width, kernel size, etc
  • LayerDrop: Introduce layer-wise dropout during training for on-demand layer reduction at test-time

47

Cai, H., Gan, C., Wang, T., Zhang, Z., & Han, S. (2019). Once-for-all: Train one network and specialize it for efficient deployment. arXiv preprint arXiv:1908.09791.

Fan, A., Grave, E., & Joulin, A. (2019). Reducing transformer depth on demand with structured dropout. arXiv preprint arXiv:1909.11556.

48 of 135

Proposed Solution

48

49 of 135

ADMN Main Contributions

  • We present the first multimodal network that allocates resources among modalities according to each modality’s quality-of-information (QoI)
  • We create a general framework for training layer-adaptive multimodal networks that can respond to dynamic computational resource availability
  • We design a multimodal perceptual controller that is trained end-to-end without reinforcement learning

49

50 of 135

ADMN Overall Architecture

50

51 of 135

ADMN Overall Architecture

51

52 of 135

Controller Architecture

52

53 of 135

QoI Supervision

  • Without adding direct/indirect incentives to attend to QoI, the network will ignore input QoI (due to sparse reward)
  • Direct QoI Supervision: Use MLP to predict corruption
  • Autoencoder Pretraining: With no QoI labels, train an autoencoder

53

54 of 135

Results

54

55 of 135

Datasets and Corruptions

55

Dataset

Attributes

GDTM

RGB, Depth Localization

MM-Fi

RGB, Depth HAR

AVE

Audio, RGB Event videos

Corruptions

Modalities/Datasets

Gaussian Noise

RGB, Depth in GDTM, MMFI

Lowlight

RGB in GDTM, AVE

Blur

RGB in GDTM

Background Audio

Audio in AVE

Rain

RGB in AVE

56 of 135

Baselines

56

Method

Explanation

Upper Bound

Allocate all layers (i.e., 12) to each backbone

Naïve Allocation

X Modality Only

Naïve Scratch

Modality Network Selection (MNS)

Train several expert models for every budget, and use a controller to select among them

Allocation Baselines: Upper Bound, Naïve Allocation, X Modality Only

New Training Baselines: Naïve Scratch, MNS

57 of 135

Main Results – GDTM Localization

57

Upper Bound utilizes 24 layers (12 for each modality backbone)

58 of 135

Main Results – Classification

58

59 of 135

ADMN Overhead

59

GDTM:

MM-Fi:

60 of 135

Discussions and Future Work

  • Summary: ADMN’s external controller allocates resources across modalities while adhering to the current budget of available resources
  • Future Work:
    • Move towards multi-instance tasks
    • Consider more runtime variations (compounding difficulty)
    • Address real-world implementation challenges

60

61 of 135

SWAN: World-Aware Adaptive Multimodal Networks for Runtime Variations

61

ECCV 2026

Wu, J., Jin, S.S., Yuan, Y., Wigness, M., Kaplan, L.M., Qiu, H., Srivastava, M.: SWAN: World-Aware Adaptive Multimodal Networks for Runtime Variations. In: Proceedings of the 19th European Conference on Computer Vision (ECCV). (2026)

62 of 135

Motivation

  • Autonomous vehicles offer complex multimodal, muti-object detection
  • Countless runtime variations:
    • Variable modality quality (weather)
    • Platform dynamics (thermals, limited energy)
    • Variable Sample Complexity (more/less targets -> harder/easer samples )
  • How to design a model compatible with real-world constraints?

62

63 of 135

Related Work

  • Multimodal 3D Detection Datasets:
    • nuScenes, Waymo Open, KITTI
    • LiDAR, camera, and sometimes radar
  • Multimodal 3D Detection Neural Networks:
    • BEVFusion: Fuse camera and LiDAR features in the Bird’s Eye View Space
    • CMT: Learned fusion of camera and LiDAR features with transformer decoder

63

Existing Multimodal AV Deep Learning Networks neglect to consider adaptation to runtime variations

Example from nuScenes

64 of 135

Variable Sample Complexity

  • Sample complexity highly diverse in AV settings:
    • Number of localization targets
    • High degrees of modality corruption (e.g., heavy fog)
  • ADMN will fully consume the entire budget regardless of sample complexity
    • No mechanism for scaling computation within a particular budget
    • Results in inefficient usage of resources (energy, compute time)

64

65 of 135

Challenges

  • How to adapt to three different runtime variations in a single multimodal network?
    • Variable Modality Quality
    • Dynamic Computational Resource Availability
    • Variable Sample Complexity
  • How to design a system compatible with real-world deployment?
    • TensorRT runtime engine
    • Different datasets

65

66 of 135

SWAN Contributions

  • First AV multimodal network modulating neural network usage according to relative QoI, sample complexity, and platform dynamics
  • Improves upon the ADMN controller with NeuralSort gradient estimation
  • Optimizes within a particular budget with two modules:
    • Adaptive SkipGate conditionally executes controller selected layers
    • Token Pruning to eliminate background tokens
  • Evaluated on real-world AV datasets on edge hardware with TensorRT runtime engine

66

67 of 135

SWAN Design

67

68 of 135

SWAN Architecture

68

69 of 135

SWAN Controller

69

 

70 of 135

NeuralSort Example

70

 

 

 

 

 

 

NeuralSort

Exploration

Exploitation

71 of 135

SWAN SkipGate Module

  •  

71

Jointly tackling several variations requires careful module design

72 of 135

SWAN Token Pruning

  • Prune tokens prior to DETR head
  • Pass features through a convolutional network with padding to generate a mask
  • Add the sum of the mask to the loss to encourage token dropping

72

73 of 135

Results

73

74 of 135

Datasets and Baselines

  • Evaluated on the MultiCorrupt variant of nuScenes
  • Benchmarked across various budgets vs. Naïve baselines and ADMN

74

MultiCorrupt darkness

75 of 135

Main Results

  • SWAN’s controller outperforms ADMN
  • Token pruning and SkipGate reduce FLOPs and Latency

75

NDS and mAP values on Multicorrupt nuScenes, C: Controller, S: SkipGate, P: Pruning

Evaluated on RTX 4090

76 of 135

Layer Allocations

76

Layer allocations across modality for controller and SkipGate on MultiCorrupt nuScenes

77 of 135

Practical Considerations

  • Does the system still perform well on a resource limited edge device?
  • How does the system generalize to new datasets/tasks?
  • Is the architecture compatible with runtime engines? How does the performance change?

77

78 of 135

Jetson Orin AGX

78

SWAN’s savings are more pronounced on weaker hardware

79 of 135

TensorRT Deployment

79

SWAN is more effective on edge hardware and with TensorRT runtime!

Localization Error (cm) on the GDTM dataset

80 of 135

Inferring Available Compute

  • Infer system state from three hardware metrics:
    • GR3D_Freq
    • GPU Active Cycles
    • Power Consumed
  • Profiling Stage: Run the TRT model with different budgets while subjecting model to various background GPU workloads (modeling multi-tenancy)
  • Training Stage: With the collected dataset, train a small MLP to predict model latency for a given layer count and hardware state

80

81 of 135

Proof-of-Concept System (Redraw Fig)

81

TensorRT model on Jetson Orin AGX; Background GPU process alternating HIGH (1900) and LOW (2); 30 ms latency threshold

  • For a given platform state, select the largest possible budget that fits inside the latency bound (30 ms) using the MLP

  • SWAN detects and responds to platform changes

82 of 135

Conclusion and Discussion

  • SWAN greatly expands upon ADMN: more complex task, NeuralSort gradient propagation, SkipGate, token pruning, and real-world deployment
  • SWAN’s components are more effective with TensorRT and real edge hardware, making it a great fit for actual usage
  • Future work can expand on the real-world systems aspect:
    • Temporal Consistency
    • Multi-tenant background processes (e.g., memory intensive, CPU intensive)

82

83 of 135

Decoupling Vision and Language: Codebook Anchored Visual Adaptation

83

CVPR 2026

J. Wu, T. Zhao, C. Liu, J. Cai, Z. Zhang, Z. Li, A. Singh, X. Xu, M. Srivastava, and J. Wu, "Decoupling Vision and Language: Codebook Anchored Visual Adaptation," in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026.

84 of 135

Background

  • Multimodal foundation models contain world knowledge and the ability to reason over novel inputs in a zero-shot manner
  • The most common multimodal foundation model is the Large Vision-Language Model (LVLM), used for Visual Question Answering (VQA)

84

85 of 135

Background - LVLMs

85

LVLMs “glue” a pretrained vision encoder and LLM together

86 of 135

Background – Discrete LVLMs

  • LVLMs cannot do image generation (lack of image token vocabulary)
  • We can discretize the output of the image encoder for unified multimodal understanding and generation

86

Vision

Encoder

Codebook

LLM

Tabby

Cat

Tell me the species of this cat and give it a tie

Forward pass of Discrete LVLM

87 of 135

Challenges

  • Current LVLMs can struggle on domain-specific tasks underrepresented in the training dataset – deployment time variation
  • Particularly, they struggle on tasks requiring expert vision signals

87

VILA-U 7B output on Plant Disease Identification (PlantVillage)

88 of 135

Existing Solutions

  • Solutions: Current works finetune the projector [1], LLM [2], or insert LoRA layers into the LLM [3] to perform domain adaptation
  • Challenges:
    • Cannot address visual side errors
    • Fine-tuning the LLM is risky – overfitting to dataset, model collapse
    • Computationally heavy – LVLMs are large (billions or trillions of parameters)

88

[1] Gregor Geigle, Radu Timofte, and Goran Glavaš. African or european swallow? benchmarking large visionlanguage models for fine-grained object classification. arXiv:2406.14496, 2024

[2] Jiawei Chen, Dingkang Yang, Yue Jiang, Mingcheng Li, Jinjie Wei, Xiaolu Hou, and Lihua Zhang. Efficiency in focus: Layernorm as a catalyst for fine-tuning medical visual language models. In ACM International Conference on Multimedia, 2024.

[3] Llava-radz: Can multimodal large language models effectively tackle zero-shot radiology recognition?

Can we finetune the vision encoder without involving the massive LLM?

89 of 135

Proposed Solution

89

90 of 135

Intuition

  • Leverage unique property of discrete LVLMs – the discrete visual codebook
  • During domain-specific fine-tuning, the vision encoder performs a better selection of the frozen discrete visual embeddings

90

Vision

Encoder

Codebook

LLM

1B

Bengal

Cat

Finetuned

Vision

Encoder

Codebook

LLM

1B

Tabby

Cat

Before Fine-Tuning

After Fine-Tuning

91 of 135

Intuition

  • The improved vision encoder can be seamlessly substituted into other LVLMs in a direct manner, provided they share the same codebook

91

Substitute Vision Encoder Finetuned with 1B LLM

92 of 135

CRAFT Key Contributions

  • CRAFT (Codebook RegulAted Fine-Tuning) is a lightweight domain adaptation framework fine-tuning only the vision encoder while enabling cross-LLM compatibility
  • We propose a surrogate-based fine-tuning framework where improvements on a smaller network transfer over to a larger network
  • CRAFT improves domain-specific performance by 13.51% points across 10 benchmarks

92

93 of 135

CRAFT

93

94 of 135

Losses

  • Three losses: Surrogate Alignment Loss, Commitment Loss, Contrastive Loss

  • Surrogate SFT Loss:

  • Commitment Loss:

  • Contrastive Loss:

94

95 of 135

CRAFT: Token Pruning

  • Discrete tokens provide a unique opportunity
    • From the training set, track how often tokens occur
  • Frequently occurring codebook entries are pruned more severely
  • To decide which tokens for index k are pruned, rank by discretization residual and distance from neighbors

95

96 of 135

Results

96

97 of 135

Datasets and Baselines

  • 10 total datasets: VQA and repurposed fine-grained classification
    • IconQA, OCRVQA, ScienceQA, VQARAD
    • EuroSAT, Flowers-102, Kvasir, PlantVillage, Stanford-Cars, Stanford-Dogs
  • Two LLM families
    • Llama2-7B (used in VILA models)
    • Qwen-2/2.5 (0.5B, 1.5B, 3B)
  • All discrete LVLMs aligned to the same frozen codebook and visual encoder
  • Baselines: Vision Encoder FT, Projector FT, LDIFS (L2 distance)

97

98 of 135

Main Results

Surrogate Fine-tuning vs. Baselines averaged across 10 benchmarks

98

  • CRAFT greatly outperforms the Zero-Shot VILA-U-7B model
  • CRAFT is competitive or outperforms baselines

99 of 135

Quality of Generated Output

99

  • Do the fine-tuned LVLMs retain explanation/reasoning abilities?
  • Ask them to justify answers, score with Claude LVLM

Correctness, Presence, Relevance, Faithfulness, Verbosity of Explanations

100 of 135

Cross-Compatibility

100

Cross-LLM Transfer, accuracy averaged across 10 benchmarks

  • Surrogate fine-tuning is highly effective

101 of 135

Training Efficiency

101

Training on Stanford-Dogs Dataset for 1 Epoch

102 of 135

Discussion and Conclusions

  • CRAFT presents a lightweight and cross-compatible visual finetuning framework
  • Discrete codebook is a powerful tool for enforcing compatibility between finetuned vision encoder and frozen LLM

102

103 of 135

THESIS CONCLUSION

103

104 of 135

Thesis Conclusion

  • Central Theme: Model adaptation greatly mitigates inference-time variations
  • FlexLoc establishes the potential of model adaptation under deployment time sensor perspective shift
  • ADMN showcases how adaptation is similarly effective against runtime variations for accuracy and computation
  • SWAN extends ADMN to complex datasets (3D detection) with real-world considerations
  • CRAFT highlights the unique challenges of adapting general purpose foundation models

104

105 of 135

Future Work (1)

Runtime Adaptation of Foundation Models

  • Multimodal foundations models should also adapt to runtime variations
    • Architectural depth, Token Length (pruning), Quantization, Reasoning Token Effort
  • As shown in SWAN, usage of several adaptation modules requires careful coordination
  • Which of these adaptation axes is more effective?

105

106 of 135

Future Work (2)

Physical Sensor Adaptation

  • Adapt both the model and the physical sensors capturing input
    • In a dark environment, change camera shutter speed, ISO, aperture, etc
    • Train a policy network to decide optimal capture parameters for a downstream model
  • New challenges in the multimodal setting:
    • Asymmetric power consumption across sensors
    • One sensor’s output can be used to select the optimal configurations of another (e.g., depth informing camera focus)

106

107 of 135

Acknowledgements

Advisor: Mani Srivastava

Funding Agencies: CONIX, IoBT, NDSEG

NESL Collaborators: Ziqi Wang, Xiaomin Ouyang, Yuyang Yuan, Ankur Sarker, Lucas Jeong

External Collaborators: Lance Kaplan (ARL), Benjamin Marlin (UMass), Colin Samplawski (UMass)

107

108 of 135

THANK YOU

108

109 of 135

Backup – Alternative Multimodal Structures

109

110 of 135

Backup – Conditional Batch Normalization

110

Modulating early visual processing by language

Neurips 2017

111 of 135

Backup – Gumbel Softmax

111

112 of 135

Backup – ADMN Corruptions

112

  • Gaussian Noise: We add N(0, σij ) to each modality i’s input data. Each modality defines a set of Ni standard deviations {σi1 , σi2 , ...σiNi } from which σij is drawn for each sample. This setting can represent systems with unstable links injecting different levels of noise, or sensors with different settings such as camera ISO levels. We apply this to the RGB and depth modalities of the GDTM and MM-Fi datasets with Ni = 4.
  • Rain: For the outdoor RGB samples of the AVE Dataset, we also explore simulating rainfall and haze corruption, following the technique in [31].
  • Lowlight: We mimic lowlight corruption through a combination of gamma correction, color shift, and additive noise [32]. We create moderately and severely impaired lowlight RGB samples from GDTM while leaving the IR depth unchanged. However, under normal lighting (RGB unchanged), we emulate a light-saturated IR depth sensor by utilizing a frame of sunlight saturated IR depth. We also add lowlight corruption to both rainy outdoor and clean indoor RGB samples of the AVE Dataset.
  • Blur: We employ Gaussian Kernel blurring with two kernel sizes to emulate different severity of blur on RGB images from the GDTM.
  • Background Noise: We mix in wind audio for outdoor events and sound from a standing fan for indoor events to corrupt the AVE dataset’s audio modality.

113 of 135

Backup – ADMN Three Modalities

113

114 of 135

Backup – ADMN Train from Scratch

114

115 of 135

Backup – ADMN Visual Results

115

116 of 135

Backup – ADMN Visual Results

116

117 of 135

Backup – ADMN Universal Controller

  • Currently, ADMN leverages an independent controller for each layer allocation
    • Controllers are easy to train (e.g., 30 minutes) and lightweight (2.3% of model parameters)
  • Explore a universal controller where the budget information is encoded as a learnable token

117

118 of 135

SWAN vs ADMN

118

119 of 135

SWAN Visual Example

119

120 of 135

SWAN Pruning

120

121 of 135

SWAN Latency Based

121

122 of 135

SWAN More Datasets

122

123 of 135

CRAFT Full Table 1

123

124 of 135

CRAFT Language

124

125 of 135

CRAFT Loss Ablation

125

126 of 135

CRAFT Explanations

126

127 of 135

CRAFT Cross Compat

127

128 of 135

CRAFT Discrete Indices

128

129 of 135

Figures/Content

129

Wu, J., Yuan, Y., Yang, K., Kaplan, L., & Srivastava, M. (2025). ADMN: A Layer-Wise Adaptive Multimodal Network for Dynamic Input Noise and Compute Resources. Advances in Neural Information Processing Systems

130 of 135

130

Multimodal Network

Multimodal Inputs

World Info

Sensor Pose

Modality Quality

Platform Status

New Data

Adapt the weights or architecture of the multimodal neural network according to information collected about the world to mitigate inference-time variations

Key Idea

131 of 135

131

Task Specific Models

Foundation Models

Localization

Deployment Time

Sensor Perspective Shift

Localization

Classification

Runtime

Varying Modality Quality; Compute Resources

3D Detection in AVs

Runtime

Varying Modality Quality; Compute Resources; Sample Complexity

VQA

Deployment Time

Highly Specialized Domains

132 of 135

132

Task Specific Models

Foundation Models

Localization

Deployment Time

Sensor Perspective Shift

Localization

Classification

Runtime

Varying Modality Quality; Compute Resources

3D Detection in AVs

Runtime

Varying Modality Quality; Compute Resources; Sample Complexity

VQA

Deployment Time

Highly Specialized Domains

133 of 135

133

Task Specific Models

Foundation Models

Localization

Deployment Time

Sensor Perspective Shift

Localization

Classification

Runtime

Varying Modality Quality; Compute Resources

3D Detection in AVs

Runtime

Varying Modality Quality; Compute Resources; Sample Complexity

VQA

Deployment Time

Highly Specialized Domains

134 of 135

Problem Space

134

Task Specific Models

Foundation Models

Localization

Deployment Time

Sensor Perspective Shift

Localization

Classification

Runtime

Varying Modality Quality; Compute Resources

3D Detection in AVs

Runtime

Varying Modality Quality; Compute Resources; Sample Complexity

VQA

Deployment Time

Highly Specialized Domains

135 of 135

135