Deep Learning · Multimodal Fusion

Multimodal Sentiment Classification

One model, three senses. Text, image, and audio are encoded by RoBERTa, ViT, and wav2vec2, fused into a single representation, and classified into negative, neutral, or positive sentiment — trained on the MSCTD dialogue dataset.

Architecture

Late fusion of frozen-vocabulary transformer encoders with a compact MLP head.

Text
RoBERTa-large
Mean-pooled sentence embedding, 1024-d
Vision
DINOv2-large
Scene embedding from the frame, 1024-d
Audio
wav2vec2-base
Mean-pooled waveform features, optional
Fusion
Concat → MLP
Negative · Neutral · Positive

Playground

Runs against your own inference server — nothing is uploaded anywhere else. Start it locally, then analyze anything.

This is the best day of my whole life! The package arrived broken and nobody answers my emails. The meeting was moved to Thursday afternoon.
Image — upload or use your camera
preview
Audio — upload or record your voice