A comprehensive framework designed to cultivate VLMs with human-like visuospatial abilities.
Ray Yang
rayruiyang
AI & ML interests
None yet
Recent Activity
upvoted a paper about 5 hours ago
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone upvoted a paper 14 days ago
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation updated a collection about 1 month ago
VSTOrganizations
None yet