Hi there, I am a Postdoctoral Scholar at
Stanford University, working with Professors Sean Follmer and Hari Subramonyam. I am completing my Ph.D. at the
Human-Computer Interaction Institute at
Carnegie Mellon University's School of Computer Science, advised by Professor Nik Martelaro. I received my Bachelor's in Computer Science from
HKUST (full scholarship), advised by Professors Xiaojuan Ma and Kwang-Ting Cheng. Previously, I worked as a research intern at Adobe Research and Runway ML. My research is supported by Google, the Toyota Research Institute, Adobe, and Accenture.
Research Interests
My research vision is to make AI a co-creative partner for designers.
I build tools that let designers steer AI through intuitive and controllable representations: sketching with generative scaffolds, assembling model puzzle pieces, and exploring latent space maps. I also work on AI for video, including adding sound effects, detecting highlights, and finding match-cut transitions.
My work sits at the intersection of Human-Computer Interaction, Computer Vision, and Design.
Publications
Visual Lyrics: Generating Animated Text for Music Lyric Videos with an Augmented Text Editor
ACM Conference on Intelligent User Interfaces (IUI), 2026
Visual Lyrics generates animated text for music lyric videos with an augmented text editor. It combines music analysis and LLM-generated animation code, with a public dataset of over 300 creative text animations.
Inkspire: Supporting Design Exploration with Generative AI through Analogical Sketching
ACM Conference on Human Factors in Computing Systems (CHI), 2025
Inkspire helps product designers explore ideas with AI-assisted sketching, analogical inspiration, and a sketch-to-design-to-sketch feedback loop.
Jigsaw: Supporting Designers to Prototype Multimodal Applications by Chaining AI Foundation Models
ACM Conference on Human Factors in Computing Systems (CHI), 2024
Jigsaw lets designers build creative AI workflows by combining models for different tasks and media with compatible puzzle pieces.
VideoMap: Supporting Video Editing Exploration, Brainstorming, and Prototyping in the Latent Space
ACM Creativity and Cognition (C&C), 2024
NeurIPS Machine Learning for Creativity and Design, 2022
VideoMap helps video editors organize footage, find transitions, and prototype rough cuts by exploring video frames on a visual map.

Tracing Creativity: A Design Space For Creative Activity Traces in HCI
ACM Conference on Human Factors in Computing Systems (CHI), 2026
We reviewed 133 creativity systems to map how creator activity traces are captured and used, providing a design space for leveraging trace data in future creativity tools.

BioSpark: Beyond Analogical Inspiration to LLM-augmented Transfer
ACM Conference on Human Factors in Computing Systems (CHI), 2025
We developed an interactive system that helps designers discover analogical biology inspirations and transfer them to target domains.

NoTeeline: Supporting Real-Time, Personalized Notetaking with LLM-Enhanced Micronotes
ACM Conference on Intelligent User Interfaces (IUI), 2025
We built an interactive notetaking tool that lets users write quick keypoints while watching educational videos then automatically expands them into full notes.
Gen4Gen: Generative Data Pipeline for Generative Multi-Concept Composition
British Machine Vision Conference (BMVC), 2025
We developed a pipeline and dataset for benchmarking multi-concept personalized text-to-image diffusion models.
Videogenic: Identifying Highlight Moments in Videos with Professional Photographs as a Prior
ACM Creativity and Cognition (C&C), 2024
NeurIPS Machine Learning for Creativity and Design, 2022
Videogenic finds video highlights using photographs as examples of moments to look for. Editors can use a photo collection to create highlight videos for different activities.
Soundify: Matching Sound Effects to Video
ACM Symposium on User Interface Software and Technology (UIST), 2023
NeurIPS Machine Learning for Creativity and Design, 2021
Soundify matches sound effects to video, synchronizes them with the action, and adjusts panning and volume to create spatial audio.
Learning Personal Style from Few Examples
ACM Conference on Designing Interactive Systems (DIS), 2021
PseudoClient learns personal graphic design preferences from a handful of examples to help designers understand a client’s visual style.
ARchitect: Building Interactive Virtual Experiences from Physical Affordances by Bringing Human-in-the-Loop
ACM Conference on Human Factors in Computing Systems (CHI), 2020
ARchitect lets an assistant use augmented reality to map physical objects to virtual objects with matching interactions, so a VR player can use their surroundings.
SeqDynamics: Visual Analytics for Evaluating Online Problem-solving Dynamics
Eurographics Conference on Visualization (EuroVis), 2020
We developed an interactive visual analytics system for instructors to evaluate problem-solving dynamics of student learners.
Learning to Film from Professional Human Motion Videos
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019
We developed an automatic drone cinematography system by learning from cinematic drone videos captured by professionals.