Kohsuke Ide

I'm a Research Assistant at AIST and a Master's student in Computer Science at University of Tsukuba, where I work at Satoh Lab. My research focuses on visual world models, 3D understanding, vision-language models, and diagnostic evaluation.

I received my B.Sc. in Applied Mechanics and Aerospace Engineering (Minor: Computer Science) from Waseda University. Previously, I worked at Preferred Networks on 3D reconstruction and free-viewpoint video technologies.

Email  /  CV  /  Scholar  /  LinkedIn  /  Github  /  X

profile photo

Research

I'm interested in computer vision, 3D understanding, and vision-language models. My research focuses on learning 3D representations and understanding spatial relationships using large language models. Some papers are highlighted.

Three controlled diagnostics for point-cloud self-supervised learning: pretext dependency, support reliance, and score formation Beyond Downstream Scores: Controlled Diagnostics for Point-Cloud Self-Supervised Learning Evaluation
Kohsuke Ide, Ryousuke Yamada, Yue Qiu, Yoshihiro Fukuhara, Hirokatsu Kataoka, Yuki M. Asano, Yutaka Satoh
Conference on Neural Information Processing Systems (NeurIPS), 2026 (E&D Track)
Spotlight (top 1.28%; 48 oral/spotlight papers out of 3,757 submissions)

Controlled diagnostics for identifying what point-cloud self-supervised learning methods learn beyond downstream benchmark scores.

Beyond Single Object: Learning 3D Relations with Large Language Models
Kohsuke Ide, Ryousuke Yamada, Yue Qiu, Xianzheng Ma, Yoshihiro Fukuhara, Hirokatsu Kataoka, Yutaka Satoh
The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings, 2026
project page / arXiv / code / press release (jp) / press release (en)

Investigating how large language models can learn and reason about 3D spatial relationships between multiple objects.

3D sans 3D Scans: Scalable Pre-training from Video-Generated Point Clouds
Ryousuke Yamada, Kohsuke Ide, Yoshihiro Fukuhara, Hirokatsu Kataoka, Gilles Puy, Andrei Bursuc, Yuki M. Asano
The IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
project page / arXiv / code

Scalable 3D pre-training approach that generates point clouds from videos, eliminating the need for expensive 3D scans.

VGI White Paper
Visual General Intelligence: A White Paper
Hirokatsu Kataoka, Yoshihiro Fukuhara, Yonglong Tian, Shangzhe Wu, Oishi Deb, Ryousuke Yamada, Christian Rupprecht, Jianyuan Wang, Kohsuke Ide, Koichi Namekata, Xianzheng Ma, Yiming Chen, Robert Geirhos, Aditi Raghunathan, Yuki M. Asano, Deva Ramanan, David Fouhey, Andrew J. Davison, Yilun Du, Jiajun Wu, Zhuang Liu
Preprint, 2026
arXiv

Reconsidering intelligence from a vision-centered perspective and outlining research directions toward visual general intelligence.

Seeing Red, Thinking Bad: Color Bias in Vision Language Models
Kohsuke Ide, Ryousuke Yamada, Yoshihiro Fukuhara, Hirokatsu Kataoka, Yutaka Satoh
International Conference on Pattern Recognition (ICPR), 2026
project page / arXiv / blog / code

Investigating how color biases in vision language models affect their reasoning and decision-making.

Colors You Can't See, Semantic Biases You Can't Ignore
Kohsuke Ide, Ryousuke Yamada, Yoshihiro Fukuhara, Hirokatsu Kataoka, Yutaka Satoh
The IEEE/CVF International Conference on Computer Vision (ICCV) Workshop on MMRAgI, 2025

Studying the sensitivity of vision-language models to visual styling of text.

Invited Talks

Seeing What Matters, Learning How to Find Out: Toward Visual General Intelligence
CMU Video Model Journal Club (online), September 18, 2026 (PT)
Recording

Blog

Intelligence Learns How to Find Out
September 2026
Perception, action, and the making of evidence.

Plato Is Not a Space
July 2026
Predictive interfaces, not universal embeddings: from PRH to video-text alignment, event semantics, and world models.

Where Do Good Vision Targets Come From?
June 2026
Notes in the margins of James Chen's pure-vision scaling post, after the CVPR Bitter Lessons workshop: the search for a vision scaling law is a search through targets, not losses.

Color as a Presentation Variable
April 2026
An interpretive companion to “Seeing Red, Thinking Bad”: reading color bias in vision language models through the Platonic and Umwelt hypotheses.

Experience

National Institute of Advanced Industrial Science and Technology (AIST)
Research Assistant, Sep 2025 - Present; Research Intern, Jan 2024 - Aug 2025
Computer vision and pattern analysis research, focusing on VLMs and 3D representation.

Preferred Networks (PFN)
Research Intern, Aug 2024 - Mar 2025
3D reconstruction and free-viewpoint video technologies.

LightBlue Technology
ML Engineer Intern, Sep 2021 - Mar 2024
Developed light-weight object tracking algorithms for edge devices.

M3
MLOps Engineer Intern, Sep 2023 - Oct 2023
Developed open-source solution for automatic Kubernetes OOM recovery.

Awards & Projects

MIRU 2025 Interactive Presentation Award (Top 4%) - "Can 3D Large Language Models Count to Three?"

Winner of 2023 Hackathon at WINC - Pushups Counter: Work out helper using computer vision for edge device.

Academic Service

Workshop Organizer - Workshop on Visual General Intelligence: Vision Research Toward the AGI Era, CVPR 2026

Organizer - LENS (Young Researchers in Vision)
A Japanese academic community of over 400 members ranging from high school students to professors.


Template from Jon Barron's website.