docs/source/en/api/pipelines/shap_e.md
The Shap-E model was proposed in Shap-E: Generating Conditional 3D Implicit Functions by Alex Nichol and Heewoo Jun from OpenAI.
The abstract from the paper is:
We present Shap-E, a conditional generative model for 3D assets. Unlike recent work on 3D generative models which produce a single output representation, Shap-E directly generates the parameters of implicit functions that can be rendered as both textured meshes and neural radiance fields. We train Shap-E in two stages: first, we train an encoder that deterministically maps 3D assets into the parameters of an implicit function; second, we train a conditional diffusion model on outputs of the encoder. When trained on a large dataset of paired 3D and text data, our resulting models are capable of generating complex and diverse 3D assets in a matter of seconds. When compared to Point-E, an explicit generative model over point clouds, Shap-E converges faster and reaches comparable or better sample quality despite modeling a higher-dimensional, multi-representation output space.
The original codebase can be found at openai/shap-e.
[!TIP] See the reuse components across pipelines section to learn how to efficiently load the same components into multiple pipelines.
Make sure you have the following libraries installed.
# uncomment to install the necessary libraries in Colab
#!pip install -q diffusers transformers accelerate trimesh
To generate a gif of a 3D object, pass a text prompt to the [ShapEPipeline]. The pipeline generates a list of image frames which are used to create the 3D object.
import torch
from diffusers import ShapEPipeline
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
pipe = ShapEPipeline.from_pretrained("openai/shap-e", torch_dtype=torch.float16, variant="fp16")
pipe = pipe.to(device)
guidance_scale = 15.0
prompt = ["A firecracker", "A birthday cupcake"]
images = pipe(
prompt,
guidance_scale=guidance_scale,
num_inference_steps=64,
frame_size=256,
).images
Now use the [~utils.export_to_gif] function to convert the list of image frames to a gif of the 3D object.
from diffusers.utils import export_to_gif
export_to_gif(images[0], "firecracker_3d.gif")
export_to_gif(images[1], "cake_3d.gif")
<figcaption class="mt-2 text-center text-sm text-gray-500">prompt = "A firecracker"</figcaption>
<figcaption class="mt-2 text-center text-sm text-gray-500">prompt = "A birthday cupcake"</figcaption>
To generate a 3D object from another image, use the [ShapEImg2ImgPipeline]. You can use an existing image or generate an entirely new one. Let's use the Kandinsky 2.1 model to generate a new image.
from diffusers import DiffusionPipeline
import torch
prior_pipeline = DiffusionPipeline.from_pretrained("kandinsky-community/kandinsky-2-1-prior", torch_dtype=torch.float16, use_safetensors=True).to("cuda")
pipeline = DiffusionPipeline.from_pretrained("kandinsky-community/kandinsky-2-1", torch_dtype=torch.float16, use_safetensors=True).to("cuda")
prompt = "A cheeseburger, white background"
image_embeds, negative_image_embeds = prior_pipeline(prompt, guidance_scale=1.0).to_tuple()
image = pipeline(
prompt,
image_embeds=image_embeds,
negative_image_embeds=negative_image_embeds,
).images[0]
image.save("burger.png")
Pass the cheeseburger to the [ShapEImg2ImgPipeline] to generate a 3D representation of it.
from PIL import Image
from diffusers import ShapEImg2ImgPipeline
from diffusers.utils import export_to_gif
pipe = ShapEImg2ImgPipeline.from_pretrained("openai/shap-e-img2img", torch_dtype=torch.float16, variant="fp16").to("cuda")
guidance_scale = 3.0
image = Image.open("burger.png").resize((256, 256))
images = pipe(
image,
guidance_scale=guidance_scale,
num_inference_steps=64,
frame_size=256,
).images
gif_path = export_to_gif(images[0], "burger_3d.gif")
<figcaption class="mt-2 text-center text-sm text-gray-500">cheeseburger</figcaption>
<figcaption class="mt-2 text-center text-sm text-gray-500">3D cheeseburger</figcaption>
Shap-E is a flexible model that can also generate textured mesh outputs to be rendered for downstream applications. In this example, you'll convert the output into a glb file because the 🤗 Datasets library supports mesh visualization of glb files which can be rendered by the Dataset viewer.
You can generate mesh outputs for both the [ShapEPipeline] and [ShapEImg2ImgPipeline] by specifying the output_type parameter as "mesh":
import torch
from diffusers import ShapEPipeline
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
pipe = ShapEPipeline.from_pretrained("openai/shap-e", torch_dtype=torch.float16, variant="fp16")
pipe = pipe.to(device)
guidance_scale = 15.0
prompt = "A birthday cupcake"
images = pipe(prompt, guidance_scale=guidance_scale, num_inference_steps=64, frame_size=256, output_type="mesh").images
Use the [~utils.export_to_ply] function to save the mesh output as a ply file:
[!TIP] You can optionally save the mesh output as an
objfile with the [~utils.export_to_obj] function. The ability to save the mesh output in a variety of formats makes it more flexible for downstream usage!
from diffusers.utils import export_to_ply
ply_path = export_to_ply(images[0], "3d_cake.ply")
print(f"Saved to folder: {ply_path}")
Then you can convert the ply file to a glb file with the trimesh library:
import trimesh
mesh = trimesh.load("3d_cake.ply")
mesh_export = mesh.export("3d_cake.glb", file_type="glb")
By default, the mesh output is focused from the bottom viewpoint but you can change the default viewpoint by applying a rotation transform:
import trimesh
import numpy as np
mesh = trimesh.load("3d_cake.ply")
rot = trimesh.transformations.rotation_matrix(-np.pi / 2, [1, 0, 0])
mesh = mesh.apply_transform(rot)
mesh_export = mesh.export("3d_cake.glb", file_type="glb")
Upload the mesh file to your dataset repository to visualize it with the Dataset viewer!
<div class="flex justify-center"> </div>[[autodoc]] ShapEPipeline - all - call
[[autodoc]] ShapEImg2ImgPipeline - all - call
[[autodoc]] pipelines.shap_e.pipeline_shap_e.ShapEPipelineOutput