When was multimodal learning proposed?
Multimodal learning was proposed in 2011, at the beginning of the deep learning period.
Short answers, pulled from the story.
Multimodal learning was proposed in 2011, at the beginning of the deep learning period.
Multimodal learning integrates modalities such as text, audio, images, and video to build a more complete understanding of complex data.
CLIP, or Contrastive Language-Image Pretraining, is a multimodal model that learns joint representations of images and text by optimizing contrastive objectives. This lets it match images with their corresponding textual descriptions.
The Boltzmann machine was invented by Geoffrey Hinton and Terry Sejnowski in 1985. It is a stochastic neural network and the generative counterpart of Hopfield nets.
Multimodal learning is applied in cross-modal retrieval, classification and missing data prediction, healthcare diagnostics such as cancer screening, content generation like DALL·E's image creation, robotics and human-computer interaction, and emotion recognition.
Google Gemini and GPT-4o became increasingly popular large multimodal models starting in 2023, enabling increased versatility and a broader understanding of real-world phenomena.