LLM-wrapper:视觉语言基础模型的黑盒语义感知适配

Amaia Cardiel, Eloi Zablocki, Oriane Siméoni, Elias Ramzi, Matthieu Cord
{"title":"LLM-wrapper:视觉语言基础模型的黑盒语义感知适配","authors":"Amaia Cardiel, Eloi Zablocki, Oriane Siméoni, Elias Ramzi, Matthieu Cord","doi":"arxiv-2409.11919","DOIUrl":null,"url":null,"abstract":"Vision Language Models (VLMs) have shown impressive performances on numerous\ntasks but their zero-shot capabilities can be limited compared to dedicated or\nfine-tuned models. Yet, fine-tuning VLMs comes with limitations as it requires\n`white-box' access to the model's architecture and weights as well as expertise\nto design the fine-tuning objectives and optimize the hyper-parameters, which\nare specific to each VLM and downstream task. In this work, we propose\nLLM-wrapper, a novel approach to adapt VLMs in a `black-box' manner by\nleveraging large language models (LLMs) so as to reason on their outputs. We\ndemonstrate the effectiveness of LLM-wrapper on Referring Expression\nComprehension (REC), a challenging open-vocabulary task that requires spatial\nand semantic reasoning. Our approach significantly boosts the performance of\noff-the-shelf models, resulting in competitive results when compared with\nclassic fine-tuning.","PeriodicalId":501130,"journal":{"name":"arXiv - CS - Computer Vision and Pattern Recognition","volume":null,"pages":null},"PeriodicalIF":0.0000,"publicationDate":"2024-09-18","publicationTypes":"Journal Article","fieldsOfStudy":null,"isOpenAccess":false,"openAccessPdf":"","citationCount":"0","resultStr":"{\"title\":\"LLM-wrapper: Black-Box Semantic-Aware Adaptation of Vision-Language Foundation Models\",\"authors\":\"Amaia Cardiel, Eloi Zablocki, Oriane Siméoni, Elias Ramzi, Matthieu Cord\",\"doi\":\"arxiv-2409.11919\",\"DOIUrl\":null,\"url\":null,\"abstract\":\"Vision Language Models (VLMs) have shown impressive performances on numerous\\ntasks but their zero-shot capabilities can be limited compared to dedicated or\\nfine-tuned models. Yet, fine-tuning VLMs comes with limitations as it requires\\n`white-box' access to the model's architecture and weights as well as expertise\\nto design the fine-tuning objectives and optimize the hyper-parameters, which\\nare specific to each VLM and downstream task. In this work, we propose\\nLLM-wrapper, a novel approach to adapt VLMs in a `black-box' manner by\\nleveraging large language models (LLMs) so as to reason on their outputs. We\\ndemonstrate the effectiveness of LLM-wrapper on Referring Expression\\nComprehension (REC), a challenging open-vocabulary task that requires spatial\\nand semantic reasoning. Our approach significantly boosts the performance of\\noff-the-shelf models, resulting in competitive results when compared with\\nclassic fine-tuning.\",\"PeriodicalId\":501130,\"journal\":{\"name\":\"arXiv - CS - Computer Vision and Pattern Recognition\",\"volume\":null,\"pages\":null},\"PeriodicalIF\":0.0000,\"publicationDate\":\"2024-09-18\",\"publicationTypes\":\"Journal Article\",\"fieldsOfStudy\":null,\"isOpenAccess\":false,\"openAccessPdf\":\"\",\"citationCount\":\"0\",\"resultStr\":null,\"platform\":\"Semanticscholar\",\"paperid\":null,\"PeriodicalName\":\"arXiv - CS - Computer Vision and Pattern Recognition\",\"FirstCategoryId\":\"1085\",\"ListUrlMain\":\"https://doi.org/arxiv-2409.11919\",\"RegionNum\":0,\"RegionCategory\":null,\"ArticlePicture\":[],\"TitleCN\":null,\"AbstractTextCN\":null,\"PMCID\":null,\"EPubDate\":\"\",\"PubModel\":\"\",\"JCR\":\"\",\"JCRName\":\"\",\"Score\":null,\"Total\":0}","platform":"Semanticscholar","paperid":null,"PeriodicalName":"arXiv - CS - Computer Vision and Pattern Recognition","FirstCategoryId":"1085","ListUrlMain":"https://doi.org/arxiv-2409.11919","RegionNum":0,"RegionCategory":null,"ArticlePicture":[],"TitleCN":null,"AbstractTextCN":null,"PMCID":null,"EPubDate":"","PubModel":"","JCR":"","JCRName":"","Score":null,"Total":0}
引用次数: 0

摘要

视觉语言模型(VLM)在大量任务中表现出令人印象深刻的性能,但与专用模型或微调模型相比,它们的零拍摄能力可能有限。然而,对 VLM 进行微调也有其局限性,因为它需要 "白盒 "访问模型的架构和权重,还需要专家来设计微调目标和优化超参数,这些都是每个 VLM 和下游任务所特有的。在这项工作中,我们提出了LLM-wrapper,这是一种以 "黑箱 "方式调整VLM的新方法,通过利用大型语言模型(LLM)来对其输出进行推理。我们演示了 LLM-wrapper 在参考表达式理解(REC)上的有效性,这是一项具有挑战性的开放词汇任务,需要空间和语义推理。我们的方法大大提高了现成模型的性能,与传统的微调方法相比,结果极具竞争力。
本文章由计算机程序翻译,如有差异,请以英文原文为准。
LLM-wrapper: Black-Box Semantic-Aware Adaptation of Vision-Language Foundation Models
Vision Language Models (VLMs) have shown impressive performances on numerous tasks but their zero-shot capabilities can be limited compared to dedicated or fine-tuned models. Yet, fine-tuning VLMs comes with limitations as it requires `white-box' access to the model's architecture and weights as well as expertise to design the fine-tuning objectives and optimize the hyper-parameters, which are specific to each VLM and downstream task. In this work, we propose LLM-wrapper, a novel approach to adapt VLMs in a `black-box' manner by leveraging large language models (LLMs) so as to reason on their outputs. We demonstrate the effectiveness of LLM-wrapper on Referring Expression Comprehension (REC), a challenging open-vocabulary task that requires spatial and semantic reasoning. Our approach significantly boosts the performance of off-the-shelf models, resulting in competitive results when compared with classic fine-tuning.
求助全文
通过发布文献求助,成功后即可免费获取论文全文。 去求助
来源期刊
自引率
0.00%
发文量
0
×
引用
GB/T 7714-2015
复制
MLA
复制
APA
复制
导出至
BibTeX EndNote RefMan NoteFirst NoteExpress
×
提示
您的信息不完整,为了账户安全,请先补充。
现在去补充
×
提示
您因"违规操作"
具体请查看互助需知
我知道了
×
提示
确定
请完成安全验证×
copy
已复制链接
快去分享给好友吧!
我知道了
右上角分享
点击右上角分享
0
联系我们:info@booksci.cn Book学术提供免费学术资源搜索服务,方便国内外学者检索中英文文献。致力于提供最便捷和优质的服务体验。 Copyright © 2023 布克学术 All rights reserved.
京ICP备2023020795号-1
ghs 京公网安备 11010802042870号
Book学术文献互助
Book学术文献互助群
群 号:481959085
Book学术官方微信