image-to-3dtrellis2mit

microsoft/TRELLIS.2-4B

huggingface.co/microsoft/TRELLIS.2-4B

1,205Likes
1,923,858Downloads
2025-12-27Updated
trellis2image-to-3denarxiv:2512.14692license:mitregion:us

Model card

TRELLIS.2: Native and Compact Structured Latents for 3D Generation

Model Name: TRELLIS.2-4B

Paper: https://arxiv.org/abs/2512.14692

Repository: https://github.com/microsoft/TRELLIS.2

Project Page: https://microsoft.github.io/trellis.2

Introduction

TRELLIS.2 is a state-of-the-art large 3D generative model designed for high-fidelity image-to-3D generation. It leverages a novel "field-free" sparse voxel structure termed O-Voxel and a large-scale flow-matching transformer (4 Billion parameters).

Unlike previous methods that rely on iso-surface fields (e.g., SDF, Flexicubes) which struggle with open surfaces or non-manifold geometry, TRELLIS can reconstruct and generate arbitrary 3D assets with complex topologies, sharp features, and full Physical-Based Rendering (PBR) materials—including transparency/translucency.

Model Details

Key Features

Inference Speed (NVIDIA H100 GPU)

| Resolution | Time | | :--- | :--- | | 512³ | ~3 seconds | | 1024³ | ~17 seconds | | 1536³ | ~60 seconds |

Requirements

Known Limitations

We are actively working on improving the model and addressing these limitations.

Usage

Note: Please refer to the official GitHub Repository for installation instructions and dependencies.

```python import os os.environ['OPENCV_IO_ENABLE_OPENEXR'] = '1' os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True" # Can save GPU memory import cv2 import imageio from PIL import Image import torch from trellis2.pipelines import Trellis2ImageTo3DPipeline from trellis2.utils import render_utils from trellis2.renderers import EnvMap import o_voxel

1. Setup Environment Map

envmap = EnvMap(torch.tensor( cv2.cvtColor(cv2.imread('assets/hdri/forest.exr', cv2.IMREAD_UNCHANGED), cv2.COLOR_BGR2RGB), dtype=torch.float32, device='cuda' ))

2. Load Pipeline

pipeline = Trellis2ImageTo3DPipeline.from_pretrained("microsoft/TRELLIS.2-4B") pipeline.cuda()

3. Load Image & Run

image = Image.open("assets/example_image/T.png") mesh = pipeline.run(image)[0] mesh.simplify(16777216) # nvdiffrast limit

4. Render Video

video = render_utils.make_pbr_vis_frames(render_utils.render_video(mesh, envmap=envmap)) imageio.mimsave("sample.mp4", video, fps=15)

5. Export to GLB

glb = o_voxel.postprocess.to_glb( vertices = mesh.vertices, faces = mesh.faces, attr_volume = mesh.attrs, coords = mesh.coords, attr_layout = mesh.layout, voxel_size = mesh.voxel_size, aabb = [[-0.5, -0.5, -0.5], [0.5, 0.5, 0.5]], decimation_target = 1000000, texture_size = 4096, remesh = True, remesh_band = 1, remesh_project = 0, verbose = True ) glb.export("sample.glb", extension_webp=True) ```

Citation

If you find this model useful for your research, please cite our work:

`` @article{ xiang2025trellis2, title={Native and Compact Structured Latents for 3D Generation}, author={Xiang, Jianfeng and Chen, Xiaoxue and Xu, Sicheng and Wang, Ruicheng and Lv, Zelong and Deng, Yu and Zhu, Hongyuan and Dong, Yue and Zhao, Hao and Yuan, Nicholas Jing and Yang, Jiaolong}, journal={Tech report}, year={2025} } ``

License

This model is released under the MIT License. The code and dataset are publicly released to facilitate reproduction and further research.

Mirrored from the Hugging Face Hub and served from the Conceptio Open Knowledge Archive. Read the original card at https://huggingface.co/microsoft/TRELLIS.2-4B.