Text-to-Image
Cosmos
Diffusers
Safetensors
cosmos3_omni
nvidia
cosmos3
vllm-omni
sglang
sglang-diffusion
image-generation
Instructions to use nvidia/Cosmos3-Super-Text2Image with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Cosmos
How to use nvidia/Cosmos3-Super-Text2Image with Cosmos:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Diffusers
How to use nvidia/Cosmos3-Super-Text2Image with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("nvidia/Cosmos3-Super-Text2Image", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
docs: scope model versions and parameter list to Cosmos3-Super-Text2Image
Browse files
README.md
CHANGED
|
@@ -32,18 +32,6 @@ This model is ready for commercial and non-commercial use.
|
|
| 32 |
**Model Developer:** NVIDIA
|
| 33 |
|
| 34 |
### Model Versions
|
| 35 |
-
- Cosmos3-Nano:
|
| 36 |
-
- Given multimodal inputs including text, images, video, audio, and action trajectories, generate coherent text, images, video, audio, and action outputs for multimodal understanding, world simulation, future prediction, action reasoning, and Physical AI applications.
|
| 37 |
-
|
| 38 |
-
- Cosmos3-Super:
|
| 39 |
-
- Given multimodal inputs including text, images, video, audio, and action trajectories, generate coherent text, images, video, audio, and action outputs for multimodal understanding, world simulation, future prediction, action reasoning, and Physical AI applications.
|
| 40 |
-
|
| 41 |
-
- Cosmos3-Nano-Policy-DROID:
|
| 42 |
-
- Given language instructions and visual observations from the DROID robot platform, generate robot action trajectories for manipulation and control tasks.
|
| 43 |
-
|
| 44 |
-
- Cosmos3-Super-Image2Video:
|
| 45 |
-
- Given one input image and text instructions, generate temporally coherent video sequences that are consistent with the provided visual content.
|
| 46 |
-
|
| 47 |
- Cosmos3-Super-Text2Image:
|
| 48 |
- Given text input, generate high-fidelity images that are consistent with the provided description.
|
| 49 |
|
|
@@ -76,10 +64,6 @@ Cosmos3 is an Omni-modal foundation model built on a Mixture-of-Transformers (Mo
|
|
| 76 |
|
| 77 |
**Number of trainable model parameters:**
|
| 78 |
|
| 79 |
-
- Cosmos3-Nano: 16B
|
| 80 |
-
- Cosmos3-Super: 64B
|
| 81 |
-
- Cosmos3-Nano-Policy-DROID: 16B
|
| 82 |
-
- Cosmos3-Super-Image2Video: 64B
|
| 83 |
- Cosmos3-Super-Text2Image: 64B
|
| 84 |
|
| 85 |
## Input/Output Specifications
|
|
|
|
| 32 |
**Model Developer:** NVIDIA
|
| 33 |
|
| 34 |
### Model Versions
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 35 |
- Cosmos3-Super-Text2Image:
|
| 36 |
- Given text input, generate high-fidelity images that are consistent with the provided description.
|
| 37 |
|
|
|
|
| 64 |
|
| 65 |
**Number of trainable model parameters:**
|
| 66 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 67 |
- Cosmos3-Super-Text2Image: 64B
|
| 68 |
|
| 69 |
## Input/Output Specifications
|