Abstract: GPT-4-Vision (GPT-4V) large multimodal models (LMMs) may be used to do zero-shot graphic layout design generation in a versatile manner. Segmentation/superpixel methods may be used to identify and mark the key regions to visually augment the image to enhance GPT-4V's spatial reasoning capability. The results demonstrate the efficacy of these visual prompting methods, showing improvement over standard GPT-4V prompting methods and also performing at par and even better, for some techniques, when compared to the LayoutDetr model.
Type:
Grant
Filed:
May 12, 2025
Date of Patent:
March 10, 2026
Assignee:
Fractal Analytics Limited
Inventors:
Kunal Singh, Mukund Khanna, Ankan Biswas, Pradeep Moturi, Shivam