DreamBeast: Distilling 3D Fantastical Animals with Part-Aware Knowledge Transfer

arXiv - EE - Image and Video Processing Pub Date : 2024-09-12 DOI:arxiv-2409.08271

Runjia Li, Junlin Han, Luke Melas-Kyriazi, Chunyi Sun, Zhaochong An, Zhongrui Gui, Shuyang Sun, Philip Torr, Tomas Jakab

{"title":"DreamBeast: Distilling 3D Fantastical Animals with Part-Aware Knowledge Transfer","authors":"Runjia Li, Junlin Han, Luke Melas-Kyriazi, Chunyi Sun, Zhaochong An, Zhongrui Gui, Shuyang Sun, Philip Torr, Tomas Jakab","doi":"arxiv-2409.08271","DOIUrl":null,"url":null,"abstract":"We present DreamBeast, a novel method based on score distillation sampling\n(SDS) for generating fantastical 3D animal assets composed of distinct parts.\nExisting SDS methods often struggle with this generation task due to a limited\nunderstanding of part-level semantics in text-to-image diffusion models. While\nrecent diffusion models, such as Stable Diffusion 3, demonstrate a better\npart-level understanding, they are prohibitively slow and exhibit other common\nproblems associated with single-view diffusion models. DreamBeast overcomes\nthis limitation through a novel part-aware knowledge transfer mechanism. For\neach generated asset, we efficiently extract part-level knowledge from the\nStable Diffusion 3 model into a 3D Part-Affinity implicit representation. This\nenables us to instantly generate Part-Affinity maps from arbitrary camera\nviews, which we then use to modulate the guidance of a multi-view diffusion\nmodel during SDS to create 3D assets of fantastical animals. DreamBeast\nsignificantly enhances the quality of generated 3D creatures with\nuser-specified part compositions while reducing computational overhead, as\ndemonstrated by extensive quantitative and qualitative evaluations.","PeriodicalId":501289,"journal":{"name":"arXiv - EE - Image and Video Processing","volume":"7 1","pages":""},"PeriodicalIF":0.0000,"publicationDate":"2024-09-12","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":null,"platform":"Semanticscholar","paperid":null,"PeriodicalName":"arXiv - EE - Image and Video Processing","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/arxiv-2409.08271","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}

引用次数: 0

Abstract

We present DreamBeast, a novel method based on score distillation sampling (SDS) for generating fantastical 3D animal assets composed of distinct parts. Existing SDS methods often struggle with this generation task due to a limited understanding of part-level semantics in text-to-image diffusion models. While recent diffusion models, such as Stable Diffusion 3, demonstrate a better part-level understanding, they are prohibitively slow and exhibit other common problems associated with single-view diffusion models. DreamBeast overcomes this limitation through a novel part-aware knowledge transfer mechanism. For each generated asset, we efficiently extract part-level knowledge from the Stable Diffusion 3 model into a 3D Part-Affinity implicit representation. This enables us to instantly generate Part-Affinity maps from arbitrary camera views, which we then use to modulate the guidance of a multi-view diffusion model during SDS to create 3D assets of fantastical animals. DreamBeast significantly enhances the quality of generated 3D creatures with user-specified part compositions while reducing computational overhead, as demonstrated by extensive quantitative and qualitative evaluations.

查看原文本刊更多论文

梦幻野兽利用部分感知知识转移提炼 3D 梦幻动物

由于对文本到图像扩散模型中部件级语义的理解有限，现有的 SDS 方法往往难以完成这一生成任务。虽然新近的扩散模型（如稳定扩散 3）展示了较好的部件级理解，但它们的速度太慢，并表现出与单视角扩散模型相关的其他常见问题。DreamBeast 通过一种新颖的部分感知知识转移机制克服了这一局限。对于每个生成的资产，我们都能高效地从稳定扩散 3 模型中提取部件级知识，并将其转化为三维部件-亲和性隐式表示。这样，我们就能从任意的摄像机视图中即时生成 "部分-亲和性 "图，然后在 SDS 过程中用它来调节多视图扩散模型的引导，从而创建出奇幻动物的三维资产。正如大量定量和定性评估所证明的那样，DreamBeasts 显著提高了根据用户指定的部件组成生成的三维动物的质量，同时降低了计算开销。

本文章由计算机程序翻译，如有差异，请以英文原文为准。

求助全文

约1分钟内获得全文求助全文

来源期刊

arXiv - EE - Image and Video Processing

自引率

0.00%

发文量