Spring til hovednavigation Spring til søgning Spring til hovedindhold

Personalizing Text-to-Image Diffusion Models by Fine-Tuning Classification for AI Applications

  • Rafael Hidalgo
  • , Nesreen Salah
  • , Rajiv Chandra Jetty
  • , Anupama Jetty
  • , Aparna S. Varde*
  • *Kontaktforfatter
  • Montclair State University

Publikation: Kapitel i bog/rapport/konference-proceedingKonferencebidrag i proceedingsForskningpeer review

Abstract

Stable Diffusion is a captivating text-to-image model that generates images based on text input. However, a major challenge is that it is pretrained on a specific dataset, limiting its ability to generate images outside of the given data. In this paper, we propose to harness two models based on neural networks, Hypernetworks and DreamBooth, to allow the introduction of any image into Stable Diffusion, addressing versatility with minimal additional training data. This work targets AI applications such as augmenting next-generation multipurpose robots, enhancing human-robot collaboration, feeding intelligent tutoring systems, training autonomous cars, injecting subjects for photo personalization, producing high quality movie animations etc. It can contribute to AI in smart cities: facets such as smart living and smart mobility.

OriginalsprogEngelsk
TitelIntelligent Systems and Applications - Proceedings of the 2023 Intelligent Systems Conference IntelliSys Volume 1
RedaktørerKohei Arai
Antal sider17
ForlagSpringer Science+Business Media
Publikationsdato2024
Sider642-658
ISBN (Trykt)9783031477201
DOI
StatusUdgivet - 2024
Udgivet eksterntJa
BegivenhedIntelligent Systems Conference, IntelliSys 2023 - Amsterdam, Holland
Varighed: 7. sep. 20238. sep. 2023

Konference

KonferenceIntelligent Systems Conference, IntelliSys 2023
Land/OmrådeHolland
ByAmsterdam
Periode07/09/202308/09/2023
NavnLecture Notes in Networks and Systems
Vol/bind822
ISSN2367-3370

Bibliografisk note

Publisher Copyright:
© 2024, The Author(s), under exclusive license to Springer Nature Switzerland AG.

Finansiering

and Disclaimer Dr. Aparna Varde acknowledges NSF grants 2018575 “MRI: Acquisition of a High-Performance GPU Cluster for Research & Education”, and 2117308 “MRI: Acquisition of a Multimodal Collaborative Robot System (MCROS) to Support Cross-Disciplinary Human-Centered Research & Education at Montclair State University”. She is a visiting researcher at Max Planck Institute for Informatics, Germany (ongoing from sabbatical). She is an Associate Director of the School of Computing, and an Associate Director of the CESAC: Clean Energy & Sustainability Analytics Center, Montclair State University. We make a disclaimer that the opinions presented here are extracted from online sources; and the content of this paper including the images is not meant to offend/hurt any national, ethnic, cultural, racial, religious and other groups. The images produced here are taken with the consent of the respective subjects. Any resemblance to anyone else is coincidental. This is a pilot study. Acknowledgments. Acknowledgments and Disclaimer Dr. Aparna Varde acknowledges NSF grants 2018575 “MRI: Acquisition of a High-Performance GPU Cluster for Research & Education”, and 2117308 “MRI: Acquisition of a Multimodal Collaborative Robot System (MCROS) to Support Cross-Disciplinary Human-Centered Research & Education at Montclair State University”. She is a visiting researcher at Max Planck Institute for Informatics, Germany (ongoing from sabbatical). She is an Associate Director of the School of Computing, and an Associate Director of the CESAC: Clean Energy & Sustainability Analytics Center, Montclair State University. We make a disclaimer that the opinions presented here are extracted from online sources; and the content of this paper including the images is not meant to offend/hurt any national, ethnic, cultural, racial, religious and other groups. The images produced here are taken with the consent of the respective subjects. Any resemblance to anyone else is coincidental. This is a pilot study.

Fingeraftryk

Dyk ned i forskningsemnerne om 'Personalizing Text-to-Image Diffusion Models by Fine-Tuning Classification for AI Applications'. Sammen danner de et unikt fingeraftryk.

Citationsformater