After entering a prompt such as, "A blue poison-dart frog sitting on a water lily," Magic3D generates a 3D mesh model, complete with colored texture, in about 40 minutes.In its academic paper, Nvidia frames Magic3D as a response to DreamFusion, a text-to-3D model that Google researchers announced in September.Similar to how DreamFusion uses a text-to-image model to generate a 2D image that then gets optimized into volumetric NeRF (Neural radiance field) data, Magic3D uses a two-stage process that takes a coarse model generated in low resolution and optimizes it to higher resolution.In 2022 alone, we've seen the emergence of capable text-to-image models such as DALL-E and Stable Diffusion and rudimentary text-to-video generators from Google and Meta.Google also debuted the aforementioned text-to-3D model DreamFusion two months ago, and since then, people have adapted similar techniques to work with as an open source model based on Stable Diffusion."