Daichi Yashima

Daichi Yashima

Ph.D. Student, Keio University
JSPS Research Fellow (DC1)

I am a Ph.D. student in Computer Science at Keio University, advised by Prof. Komei Sugiura. I am supported by the JSPS Research Fellowship for Young Scientists (DC1). I started my Ph.D. in April 2026 after completing the Master's program in one year.

My research focuses on foundation models and multimodal language understanding for embodied AI: building systems that can execute complex tasks in the physical world. I work on multimodal large language models, vision-language-action models and mobile manipulation.

News

Publications

2026

MoFlo: Language-Conditioned Flow Matching for Policy Mobilization
MoFlo: Language-Conditioned Flow Matching for Policy Mobilization
D. Yashima, Y. Takagi, K. Seno, R. Suzuki, K. Sakata, S. Kurita, and K. Sugiura
CoRL 2026 (Acceptance Rate: 33%, h5-index: 107)
Flow as Flow: Modeling Robot Velocity Fields as Probability Velocity Fields for Flow-Based Object Manipulation
Flow as Flow: Modeling Robot Velocity Fields as Probability Velocity Fields for Flow-Based Object Manipulation
K. Seno, D. Yashima, Y. Takagi, K. Tokura, and K. Sugiura
CoRL 2026 (Acceptance Rate: 33%, h5-index: 107)
RIGEL: Self-Distilled Score Adaptation for Image and Video Captioning Evaluation
RIGEL: Self-Distilled Score Adaptation for Image and Video Captioning Evaluation
S. Koyama, K. Matsuda, Y. Wada, S. Hirano, D. Yashima, and K. Sugiura
EMNLP 2026 (main) (Acceptance Rate: 15.4%, h5-index: 218)
HiFlow: Tokenization-Free Scale-Wise Autoregressive Policy Learning via Flow Matching
HiFlow: Tokenization-Free Scale-Wise Autoregressive Policy Learning via Flow Matching
D. Yashima, K. Seno, S. Kurita, Y. Oda, and K. Sugiura
IROS 2026 (Acceptance Rate: 36%, h5-index: 92)
ELSA: Acoustic Event-Level Semantic Alignment for Fine-Grained Reference-Free Text-to-Audio Evaluation
ELSA: Acoustic Event-Level Semantic Alignment for Fine-Grained Reference-Free Text-to-Audio Evaluation
S. Suzuki*, K. Tokura*, D. Yashima*, K. Amemiya*, K. Sugiura, and S. Takamichi
INTERSPEECH 2026 (h5-index: 112)
MLLM-as-a-Judge Exhibits Model Preference Bias
MLLM-as-a-Judge Exhibits Model Preference Bias
S. Koyama*, Y. Wada*, D. Yashima*, and K. Sugiura
Preprint
ABMAMBA: Multimodal Large Language Model with Aligned Hierarchical Bidirectional Scan for Efficient Video Captioning
ABMAMBA: Multimodal Large Language Model with Aligned Hierarchical Bidirectional Scan for Efficient Video Captioning
D. Yashima, S. Kurita, Y. Oda, S. Suzuki, S. Otsuki, and K. Sugiura
ICPR 2026 (h5-index: 68)
AnoleVLA: Lightweight Vision-Language-Action Model with Deep State Space Models for Mobile Manipulation
AnoleVLA: Lightweight Vision-Language-Action Model with Deep State Space Models for Mobile Manipulation
Y. Takagi, M. Kambara, D. Yashima, K. Seno, K. Tokura, and K. Sugiura
Preprint
ReMoRa: Multimodal Large Language Model based on Refined Motion Representation for Long-Video Understanding
ReMoRa: Multimodal Large Language Model based on Refined Motion Representation for Long-Video Understanding
D. Yashima, S. Kurita, Y. Oda, and K. Sugiura
CVPR 2026 (Acceptance Rate: 25.42%, h5-index: 450)
NaiLIA: Multimodal Nail Design Retrieval Based on Dense Intent Descriptions and Palette Queries
NaiLIA: Multimodal Nail Design Retrieval Based on Dense Intent Descriptions and Palette Queries
K. Amemiya, D. Yashima, K. Katsumata, T. Komatsu, R. Korekata, S. Otsuki, and K. Sugiura
CVPR 2026 Findings (Acceptance Rate (main + findings): 36%, h5-index: 450)

2025

AIRoA MoMa Dataset: A Large-Scale Hierarchical Dataset for Mobile Manipulation
AIRoA MoMa Dataset: A Large-Scale Hierarchical Dataset for Mobile Manipulation
R. Takanami, P. Khrapchenkov, S. Morikuni, J. Arima, Y. Takaba, S. Maeda, T. Okubo, G. Sano, S. Sekioka, A. Kadoya, M. Kambara, N. Nishiura, H. Suzuki, T. Yoshimoto, K. Sakamoto, S. Ono, H. Yang, D. Yashima, A. Horo, T. Motoda, K. Chiyoma, H. Ito, K. Fukuda, A. Goto, K. Morinaga, Y. Ikeda, R. Kawada, M. Yoshikawa, N. Kosuge, Y. Noguchi, K. Ota, T. Matsushima, Y. Iwasawa, Y. Matsuo, and T. Ogata
Preprint
Open-Vocabulary Mobile Manipulation Based on Double Relaxed Contrastive Learning With Dense Labeling
Open-Vocabulary Mobile Manipulation Based on Double Relaxed Contrastive Learning With Dense Labeling
D. Yashima, R. Korekata, and K. Sugiura
IEEE RA-L (IF: 5.2, h5-index: 132)
Mobile Manipulation Instruction Generation from Multiple Images with Automatic Metric Enhancement
Mobile Manipulation Instruction Generation from Multiple Images with Automatic Metric Enhancement
K. Katsumata, M. Kambara, D. Yashima, R. Korekata, and K. Sugiura
IEEE RA-L (IF: 5.2, h5-index: 132)

Experience

Research

Fellowships

Talks

Paper Reading

Academic Service

Reviewer

ARR, ACCV, CoRL, NeurIPS

Domestic Conferences

  1. K. Amemiya, D. Yashima, K. Katsumata, and K. Sugiura, “Nail Design Image Retrieval with Dense Intent Descriptions and Palette Queries”, MIRU 2026, DS-03.
  2. K. Tokura, M. Goko, K. Amemiya, D. Yashima, K. Katsumata, T. Komatsu, R. Korekata, and K. Sugiura, “Scene Text-Guided Object Retrieval and Manipulation from Free-Form Instructions”, MIRU 2026, IS2-131.
  3. K. Seno, D. Yashima, Y. Takagi, K. Tokura, and K. Sugiura, “Generating Robot Flow Based on Flow Matching for Object Manipulation”, MIRU 2026, IS1-188.
  4. Y. Takagi, M. Kambara, D. Yashima, K. Seno, K. Tokura, and K. Sugiura, “Efficient Vision-Language-Action Model Using Deep State Space Models”, MIRU 2026, IS2-182.
  5. S. Koyama, Y. Wada, D. Yashima, and K. Sugiura, “To What Extent Does MLLM-as-a-Judge Exhibit Cross-Model Preference Bias?”, MIRU 2026, OS1E-03. (peer-reviewed, acceptance rate 33.5%)
  6. 細屋達稀, 小山修生, 八島大地, 和田唯我, 杉浦孔明, “Binomial Deviance ResidualによるMLLM-as-a-Judgeのモデル選好の解析”, MIRU 2026, IS1-147.
  7. 八島大地, 栗田修平, 小田悠介, 杉浦孔明, “圧縮動画表現に基づくMLLMによる動画理解”, JSAI 2026, 1Yin-A-15.
  8. 雨宮佳音, 八島大地, 勝又圭, 杉浦孔明, “パレットクエリに基づくファッション画像のマルチモーダル検索”, JSAI 2026, 1Yin-A-14.
  9. 小山修生, 和田唯我, 八島大地, 杉浦孔明, “MLLM-as-a-Judgeにおける自己選好バイアスの軽減”, JSAI 2026, 1Yin-A-06.
  10. 後神美結, 戸倉健登, 雨宮佳音, 八島大地, 勝又圭, 今井悠人, 小松拓実, 是方諒介, 杉浦孔明, “シーンテキストを用いたマルチモーダル検索に基づく日常物体操作”, RSJ 2025, 1I2-05.
  11. 西牧宙輝, 八島大地, 戸倉健登, 杉浦孔明, “多言語シーンテキストを考慮した深層状態空間モデルに基づく実世界検索エンジン”, RSJ 2025, 1M3-04.
  12. 雨宮佳音, 小松拓実, 八島大地, 是方諒介, 勝又圭, 杉浦孔明, “NaiLIA: 多層的な依頼文に基づくネイルデザインのマルチモーダル検索”, MIRU 2025, OS2B-07. (peer-reviewed, acceptance rate 34.5%)
  13. 八島大地, 栗田修平, 小田悠介, 鈴木駿太郎, 小槻誠太郎, 杉浦孔明, “深層状態空間モデルおよび双方向スキャンに基づくMultimodal LLMによる動画像理解”, MIRU 2025, IS1-140.
  14. 戸倉健登, 後神美結, 雨宮佳音, 八島大地, 勝又圭, 今井悠人, 小松拓実, 是方諒介, 杉浦孔明, “シーンテキストを考慮したCrosslingual Visual Promptに基づくマルチモーダル検索”, MIRU 2025, IS1-141.
  15. 八島大地, 是方諒介, 杉浦孔明, “二重緩和損失を用いたマルチモーダル検索に基づく生活支援ロボットによる物体操作”, JSAI 2025, 1Win4-50.
  16. 雨宮佳音, 小松拓実, 八島大地, 是方諒介, 勝又圭, 杉浦孔明, “NaiLIA: 緩和損失に基づくネイルデザインのマルチモーダル検索”, JSAI 2025, 2Win5-55.
  17. 勝又圭, 神原元就, 八島大地, 是方諒介, 杉浦孔明, “物体操作指示文生成モデルに基づくモバイルマニピュレーションのためのデータセット拡張”, JSAI 2025, 1Win4-49.
  18. 八島大地, 是方諒介, 杉浦孔明, “Multimodal LLMと二重緩和損失に基づく実世界検索エンジン”, RSJ 2024, 3D2-07. [slides]