A comprehensive framework designed to cultivate VLMs with human-like visuospatial abilities.
Ray Yang
rayruiyang
AI & ML interests
None yet
Recent Activity
upvoted a paper 17 days ago
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation updated a collection about 2 months ago
VSTOrganizations
None yet