Model reference · open weights
4M-21_L is an open-weight language model from EPFL-VILAB. 4M-21_L (FP32) weighs 3.1 GB; the smallest configuration that runs it is RTX 3060 12 GB.
What it is
| Released by | EPFL-VILAB |
|---|---|
| Type | Language models |
| Task | Omni (any→any) |
| Parameters (lead) | 1.6B |
| Runs with | ml-4m |
| Released | 2024-06-12 |
| Popularity | 510 downloads / month |
| Weights | 3.1 GB (4M-21_L (FP32), file size) |
| Licence | Its own licence terms |
What it runs on
Weights 3.1 GB (file size) · runtime overhead from 762 MB on a small card.
How much memory each request adds is not estimated yet for this architecture — only the weights are. They need the cards below at the least, plus room for the context.
| Card | The weights alone |
|---|---|
| RTX 3060 12 GB | fits |
| RTX 4060 Ti 16 GB | fits |
| RTX 3090 24 GB | fits |
| RTX 4090 24 GB | fits |
| RTX 5090 32 GB | fits |
| L40S 48 GB | fits |
| A100 80 GB | fits |
| H100 80 GB | fits |
| RTX PRO 6000 Blackwell 96 GB | fits |
| DGX Spark (GB10) 128 GB unified | fits |
| H200 141 GB | fits |
| B200 180 GB | fits |
From the model card
A framework for training any-to-any multimodal foundation models. Scalable. Open-sourced. Across tens of modalities and tasks.
Official implementation and pre-trained models for :
4M: Massively Multimodal Masked Modeling, NeurIPS 2023 (Spotlight) David Mizrahi*, Roman Bachmann*, Oğuzhan Fatih Kar, Teresa Yeo, Mingfei Gao, Afshin Dehghan, Amir Zamir
4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities, arXiv 2024 Roman Bachmann*, Oğuzhan Fatih Kar*, David Mizrahi*, Ali Garjani, Mingfei Gao, David Griffiths, Jiaming Hu, Afshin Dehghan, Amir Zamir
4M is a framework for training "any-to-any" foundation models, using tokenization and masking to scale to many diverse modalities. Models trained using 4M can perform a wide range of vision tasks, transfer well to unseen tasks and modalities, and are flexible and steerable multimodal generative models. We are releasing code and models for "4M: Massively Multimodal Masked Modeling" (here denoted 4M-7), as well as "4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities" (here denoted 4M-21).
For install instructions, please see https://github.com/apple/ml-4m.
This model can be loaded from Hugging Face Hub as follows:
from fourm.models.fm import FM
fm = FM.from_pretrained('EPFL-VILAB/4M-21_L')
Please see README_GENERATION.md for more detailed instructions and https://github.com/apple/ml-4m for other 4M model and tokenizer checkpoints.
If you find this repository helpful, please consider citing our work:
@inproceedings{4m,
title={{4M}: Massively Multimodal Masked Modeling},
author={David Mizrahi and Roman Bachmann and O{\u{g}}uzhan Fatih Kar and Teresa Yeo and Mingfei Gao and Afshin Dehghan and Amir Zamir},
booktitle={Thirty-seventh Conference on Neural Information Processing Systems},
year={2023},
}
@article{4m21,
title={{4M-21}: An Any-to-Any Vision Model for Tens of Tasks and Modalities},
author={Roman Bachmann and O{\u{g}}uzhan Fatih Kar and David Mizrahi and Ali Garjani and Mingfei Gao and David Griffiths and Jiaming Hu and Afshin Dehghan and Amir Zamir},
journal={arXiv 2024},
year={2024},
}
The model weights in this repository are released under the Sample Code license as found in the LICENSE file.
Quoted from the model card on Hugging Face — the full card is behind the Hugging Face link above.
How it works