I²MoE: learning how modalities work together

Topics: Multimodal learning · Mixture of experts · Interpretability

A movie poster and its plot can support the same genre prediction in different ways. The poster may provide a distinctive visual clue, the plot may supply information that the image leaves out, and some evidence may become useful only when the two are considered together. A fusion model should be able to learn these different relationships and show how they contribute to its answer.

An anonymous reviewer at ICML 2025 wrote: “Provides both local and global explanations, supported by qualitative and quantitative analyses.”

In our ICML 2025 paper, I²MoE: Interpretable Multimodal Interaction-aware Mixture-of-Experts, we turn these relationships into specialized interaction experts. Each expert learns a different way in which modalities can inform a prediction, and a reweighting network decides how much to use each expert for each example. The resulting model combines improved prediction with explanations at the sample and dataset levels. For the full method and experiments, here is the link to paper.

What does each interaction mean?

Partial information decomposition, or PID, separates the information that two inputs provide about a target into four components: information unique to the first input, information unique to the second, redundant information available from either input, and synergistic information that emerges from their combination.

The movie example in our introduction illustrates all four. Distinctive visual cues in the poster help identify Horror, while the plot supplies the romantic context needed for Romance. The poster’s blurry figure and the plot’s mention of a sorcerer both suggest Fantasy, providing redundant evidence. Identifying Drama draws on the combination of clothing and facial expressions in the image with narrative details in the plot. The interaction type therefore depends on the prediction we are trying to make.

I²MoE assigns a separate expert to each of these four components. The experts share modality encoders, but each has its own fusion model and prediction head. All four experts receive both modality embeddings; their training objectives encourage them to specialize in different interactions. This design also allows us to build the experts using existing fusion backbones, such as multimodal transformers.

I²MoE uses four interaction experts and a reweighting network to combine two input modalities.

I²MoE replaces a single fusion model with experts for the two unique contributions, synergy, and redundancy. A reweighting network combines their predictions according to the input example.

How do experts learn without interaction labels?

Each expert must learn the behavior associated with its interaction type. The challenge is that training data usually provide task labels, such as movie genres, rather than labels for the interactions behind them. We obtain weak supervision by comparing each expert’s prediction on the complete input with its predictions after replacing one modality embedding with a random vector.

The required behavior depends on the expert. For the first modality’s uniqueness expert, replacing the second modality should preserve the prediction, while replacing the first should change it. The redundancy expert should preserve its prediction when either modality is replaced, because the other still supplies the shared information. The synergy expert should change its prediction under either replacement, because its information depends on the two modalities together.

For classification, we implement these relationships using a triplet-margin loss for uniqueness and cosine-similarity losses for synergy and redundancy. The complete objective combines prediction quality with interaction specialization:

\[\mathcal{L}=\mathcal{L}_{\mathrm{task}}+\lambda\,\frac{1}{E}\sum_{i=1}^{E}\mathcal{L}_{\mathrm{interaction},i}.\]

Here, $E$ is the number of interaction experts, and $\lambda$ controls the contribution of the interaction losses. The model learns the encoders, experts, and reweighting network jointly, following the training procedure and complete objective in our paper.

In Appendix A, we develop the connection between this perturbation-based training and PID. We begin with the decomposition of the inputs’ mutual information with the target into unique, redundant, and synergistic components. Under the assumption that a replaced modality contains no task-relevant information, we relate prediction from a retained modality to uniqueness, successful prediction from either partial view to redundancy, and reliance on the complete view to synergy. We describe the resulting objectives as a contrastive approximation to the constrained information projections used in PID. This connects the experts’ specialization to the information-theoretic framework while providing an end-to-end training procedure.

Following a prediction through the experts

The interaction that matters most can change from one movie to the next. The reweighting network learns this choice by producing input-dependent weights, and the final prediction is their weighted sum of the experts’ outputs:

\[\hat{y}(x)=\sum_{i=1}^{E}w_i(x)\,\hat{y}_i(x).\]

For classification, each expert provides a logit for each class. We can inspect the logit, the expert’s weight, and their product to see how that expert contributes to the final score.

The Shrek example shows this calculation for a real IMDB test sample. For the Animation label, the image-uniqueness expert and redundancy expert produce positive logits, while the synergy expert produces a negative logit. The reweighting network gives higher weights to the image-uniqueness and redundancy experts, and the final weighted score is positive. The cartoon characters in the poster make the image expert’s contribution easy to understand. The same visualization shows the contributions for Comedy, Adventure, Fantasy, and Family.

The Shrek example shows expert logits, expert weights, weighted contributions, and the poster and plot used as inputs.

For Shrek, the image-uniqueness and redundancy experts receive the largest weights and support the Animation prediction. The panels trace how each expert’s output contributes to the final scores for the movie’s five genres.

We also evaluated whether people found these weight assignments reasonable. Fifteen participants each reviewed 20 movie examples, producing 300 ratings on a five-point scale. About 71% of the released ratings were positive, meaning that the weights mostly or completely made sense to the evaluator. We describe the study in Appendix H. Our official repository provides the questionnaire and de-identified user-study results.

From one example to a dataset

The same weights provide a broader view when we examine their distributions across the test set. This shows both which interaction experts receive more weight and how much their importance changes between examples.

The five datasets exhibit different patterns in Figure 4. ADNI has relatively balanced weights with modest preferences among experts. MIMIC shows much larger variation, reflecting the model’s adaptation to differences between patient records. IMDB varies less than MIMIC, while MOSI’s weights are close to uniform across interaction types. ENRICO places greater emphasis on the screenshot modality. These patterns show how the reweighting model adapts to the data rather than imposing one fixed interaction mixture.

Distributions of interaction-expert weights across the ADNI, MIMIC, IMDB, MOSI, and ENRICO test sets.

Interaction weights reveal different dataset-level patterns. MIMIC exhibits substantial sample-to-sample variation, MOSI has nearly uniform weights, and ENRICO emphasizes the screenshot modality.

Does modeling interactions improve prediction?

We evaluated I²MoE on five datasets covering Alzheimer’s disease classification, mortality prediction, movie genres, sentiment, and interface classification. To focus the comparison on fusion, the baseline and I²MoE models use the same modality-encoder and prediction-head configurations. Results are averaged over three random seeds.

With MulT as the fusion backbone, ADNI accuracy rises from 59.57 ± 0.66% to 65.08 ± 1.52%, a gain of 5.51 percentage points, and AUROC rises from 77.21 to 81.09. IMDB micro-F1 improves from 59.68 ± 0.19 to 61.00 ± 0.44, and MOSI accuracy rises from 68.80% to 71.91%. On MIMIC, AUROC is nearly unchanged, moving from 68.79 to 68.81, while accuracy falls from 72.42% to 69.78%. We discuss the accuracy–AUROC trade-off in the context of class imbalance. The complete comparison is shown in Table 1.

Table 1 compares I²MoE-MulT with seven fusion baselines across five datasets.

I²MoE-MulT improves over MulT on ADNI, IMDB, MOSI, and ENRICO. MIMIC shows a different pattern, with nearly unchanged AUROC and lower accuracy. The table reports the mean and standard deviation over three runs.

Our experiments with other fusion backbones show similar benefits. For example, adding I²MoE to SwitchGate improves ADNI accuracy from 62.28% to 67.51% and ENRICO accuracy from 43.95% to 49.09%. This supports using interaction experts as a general fusion framework rather than tying the approach to one architecture.

Our ablation study shows why the specialization matters. Removing the interaction loss lowers ADNI accuracy from 65.08% to 58.73%. Keeping only synergy and redundancy experts lowers it further to 56.77%, highlighting the contribution of the uniqueness experts. Other variants move the interaction loss to latent embeddings, replace sample-dependent weights with global weights, or reduce the number of perturbed views. Each performs below the full model on the reported metrics.

Table 4 shows the effects of removing or modifying components of I²MoE on ADNI, MOSI, and ENRICO.

All five ablations reduce performance relative to the full model. Rows 1–5 respectively remove the interaction loss, apply it to latent embeddings, use global weights, reduce perturbed views, and remove the uniqueness experts.

Extending the interaction model

Our extension to more than two modalities uses one uniqueness expert per modality, one global synergy expert, and one global redundancy expert. This gives $n+2$ experts for $n$ modalities and avoids a combinatorial expansion over modality combinations. The current design therefore represents each modality’s unique contribution and the global interactions, rather than assigning separate experts to every possible subset.

In our conclusion, we identify two directions for future work: exploring alternative interaction losses and combining interaction experts with feature attribution methods. The latter would add a finer-grained view of which features within an expert contribute to a prediction. These directions build on the central idea of I²MoE: explicitly learning how modalities work together makes their joint prediction more accurate and easier to inspect.

Reference

Jiayi Xin, Sukwon Yun, Jie Peng, Inyoung Choi, Jenna L. Ballard, Tianlong Chen, and Qi Long. 2025. I²MoE: Interpretable Multimodal Interaction-aware Mixture-of-Experts. Proceedings of the 42nd International Conference on Machine Learning, PMLR 267, pp. 68870–68888.

The implementation is available in our official code repository.

All research articles · Publications