B.Sc. Candidate in Computer Science · Graduating July 2026

Bandits, Reinforcement Learning, and Decision-making under Uncertainty.

I am a B.Sc. candidate in Computer Science at Hanoi University of Science and Technology, graduating in July 2026. My research focuses on bandits, reinforcement learning, offline/robust decision-making, and statistical learning theory. I am first author of an ICML 2026 accepted paper on variance-driven exploration and am seeking Master’s/Ph.D. opportunities and research collaborations.

Research interests

Bandit Theory Reinforcement Learning Pure Exploration Best-Arm Identification 1-Bit Mean Estimation Offline Contextual Bandits Distributionally Robust Optimization Risk-Aware Decision Making

Publications and manuscripts

  1. Khang Luong, Nam Nguyen, Hoang Ta, Hung The Tran, Tuan Dam.
    Accepted at ICML, 2026. OpenReview
  2. Nearly Optimal Fixed-Confidence Best-Arm Identification with 1-Bit Feedback
    Khang Luong, Dinh Thai Son, Hoang Ta, Hung The Tran, Tuan Dam.
    Submission to NeurIPS, 2026.

Education

Hanoi University of Science and Technology

B.Sc. in Computer Science · GPA: 3.73/4.0

2022 – Jul. 2026

Language

English: IELTS 7.0

Experience

HUST Data Science Laboratory

Research Member. Conducted research on machine learning and decision intelligence under uncertainty.

Feb. 2025 – Present

VinDynamics

AI Intern, Robotics. Worked on perception and control pipelines for robotic grasping, hand–object interaction, and vision-based teleoperation.

Jun. 2025 – May 2026

FPT Quantum AI & Cyber Security Institute

Research Member. Conducting research on offline policy evaluation and learning for contextual bandits.

Jun. 2026 – Present

Selected projects

Variance Driven Exploration (VarDE)

HUST Data Science Laboratory

  • Derived influence-weighted variance-driven sampling rules for exploration frameworks, providing theoretical guarantees on variance decay and simple regret.
  • Applied VarDE to Best-Arm Identification, Monte Carlo Tree Search, and Best-Policy Identification.

Best-Arm Identification with 1-Bit Feedback

HUST Data Science Laboratory

  • Developed a BAI algorithm with 1-bit mean estimation, matching the worst-case lower bound for 1-bit feedback.
  • Currently extending the analysis to instance-dependent bounds and asymptotically optimal parametric algorithms.
  • Applying 1-bit mean estimation to federated learning for privacy, corruption robustness, and communication efficiency.

Offline Policy Evaluation and Learning for Contextual Bandits

FPT Quantum AI & Cyber Security Institute

  • Study distributionally robust offline contextual bandits under context shift and reward shift.
  • Develop pessimistic offline policy evaluation and learning methods based on optimal-transport DRO objectives.
  • Analyze finite-sample guarantees under limited coverage and offline-data constraints.

Vision-Based Robotic Hand Teleoperation and Autonomous Manipulation

VinDynamics

  • Developed a vision-based teleoperation system for dexterous robotic hand control.
  • Worked on perception and control pipelines for robotic grasping and hand–object interaction.
  • Integrated visual perception with control policies for autonomous object manipulation.

Contact

I am interested in conversations with faculty, postdocs, Ph.D. students, and collaborators working on machine learning theory, bandits, reinforcement learning, offline/robust decision-making, and statistical learning theory.