I am interested in taking computer vision across a wide range of applications, from cryo-EM structural analysis to sports video understanding and social robot navigation.
Ultimately, my goal is to create a society where humans and AI collaborate seamlessly in all real-world environments.
Particle picking, contamination removal, and 2D class selection are usually trained and scored in isolation. Running them as one loop, and scoring the loop by the resolution of the map it delivers, beats every picker we compare — and shows that the best 2D F1 is not the best reconstruction.
Synthetic micrographs are clean and real ones are not, and that gap is usually what breaks transfer. We turn it into supervision: an anomaly detector trained on the clean synthetic data flags real contaminants, and a hard negative loss suppresses picks that land on them. Best precision, F1, and reconstruction resolution among the methods compared in the 1-shot setting on CryoPPP.
Benchmarking Social Robot Navigation in the Real World
AI Suitcase, Keio University
In progress, 2026 –
Social navigation policies are usually compared in simulation, by success rate, path length, and intrusion time. We run several of them — a non-social baseline, teleoperation as an upper bound, a reactive planner, and a legibility-driven one — on a real navigation robot in a busy campus corridor, collect human preferences over the paths, and ask which metrics actually track those preferences. Part of that question is whether a vision-language model can stand in for a user study.
Automatic Analysis of Futsal Match Video
Keio University
In progress, 2026 –
We are building the first futsal video dataset, 100 hours of match footage from the Japanese university league. On top of it we define a task set that is both practical and new to the field: player detection and tracking, set-play clipping, substitution detection, and ball detection. The goal is a futsal-specific model, since the public models trained on soccer, basketball, and volleyball only carry part of the way.
Depth Anything V2 fine-tuned for soccer broadcast frames. Person masks from Grounded SAM 2 weight the loss on players, where the baseline produces its worst artifacts, and a field-masked smoothing step removes the curvature the reconstruction shows across a pitch that should be flat.
Vocabulary Expansion for Japanese Lip Reading
Bachelor's thesis, Keio University
Thesis, 2025
Japanese lip-reading datasets cover few words, so a model cannot be asked about the rest. I built a dataset for unseen vocabulary using GAN-based talking-face generation, which widens what the model can recognize and holds up better in noise.
Experience
Research Associate, Min Xu Lab, Carnegie Mellon University — Pittsburgh, Jun – Sep 2026
R&D Intern, Edge AI, Sony Semiconductor Solutions — Tokyo, Sep 2025
Machine Learning Intern, WALC (DMG MORI subsidiary) — Tokyo, Jan 2025 – Apr 2026
Education
M.S. in Information Science, Keio University — Kanagawa, 2025 – 2027 (expected)
Hideo Saito Lab. Joint research with the Min Xu Lab at Carnegie Mellon University.
B.Eng. in System Design Engineering, Keio University — Kanagawa, 2021 – 2025
Modeling, analysis, and synthesis of complex systems.