Computer vision, natural language processing, and multimodal learning. Current focus: vision-language-action models.
컴퓨터 비전, 자연어처리, 멀티모달 학습. 최근 관심은 vision-language-action 모델.
Joint representations across image, video, and text
이미지와 영상, 텍스트를 아우르는 공통 표현
Perception and language grounded in action. Efficient adaptation of pretrained policies
인식과 언어를 행동으로 잇는 모델. 사전학습 정책의 효율적 적응
Event localization in video, described in natural language
영상 속 사건의 검출과 자연어 기술
Compression and adaptation under limited compute and memory. Cuts across the topics above
제한된 연산과 메모리에서의 경량화와 적응. 위 주제 전반에 걸침