DDMAN1101 / TAIPEI

PhD student · NTU × Academia Sinica

Wei-HanHsu

aka Jerry

Machine learning for music: DJ systems, music generation, drum transcription.

Wei-Han Hsu DJing at a pair of CDJs and a mixer
Sep 2026 — DJustify code is out on GitHub× Jul 2026 — Separate-and-Detect accepted to ISMIR 2026× Sep 2026 — DJustify code is out on GitHub× Jul 2026 — Separate-and-Detect accepted to ISMIR 2026×

Lineup

01 · under reviewCo-first author

DJustify

Reason-and-Verify DJ Planning for Pop Transitions and Story-Driven Sets

An LLM plans each transition; a rule-based checker measures the rendered audio and sends back what went wrong.

Wei-Han Hsu*, Li-Jie Lin*, Po-Hsuan Lai, Po-Hsiang Huang, Jeng-Yue Liu, Li Su, Yi-Hsuan Yang  *equal contribution

DJustify pipeline: pre-processing, retrieval, an LLM planner, rendering, and a checker that sends failures back for repair; the same loop builds story-driven sets.
FIG — transition and set systems

02 · ISMIR 2026First author

Separate-and-Detect

Unified Drum Transcription and Stem Generation via Latent Diffusion

Split a drum kit into five stems with a latent diffusion model, then transcribe each stem on its own.

Wei-Han Hsu, Chih-Cheng Chang, Bo-Yu Chen, Li Su, Yi-Hsuan Yang

Separate-and-Detect pipeline: Demucs isolates drums, a latent diffusion demixer produces five drum stems, and a per-class onset detector transcribes them.
FIG — separate, then detect

03 · arXiv 2021First author

EDM Subgenres

Deep Learning Based EDM Subgenre Classification using Mel-Spectrogram and Tempogram Features

Adds tempogram features to a music auto-tagging CNN so it can tell EDM subgenres apart.

Wei-Han Hsu, Bo-Yu Chen, Yi-Hsuan Yang

EDM subgenre model: a short-chunk CNN on the mel spectrogram fused with a model on Fourier and autocorrelation tempograms.
FIG — mel + tempogram fusion

04 · ICASSP 2022Second author

DJtransGAN

Automatic DJ Transitions with Differentiable Audio Effects and Generative Adversarial Networks

A GAN with a differentiable DJ mixer that learns fader and EQ curves from real DJ mixes. My part was the DJ side: domain knowledge and the training data.

Bo-Yu Chen, Wei-Han Hsu, Wei-Hsiang Liao, Marco A. Martínez Ramírez, Yuki Mitsufuji, Yi-Hsuan Yang

DJtransGAN: two tracks go through STFT, an encoder-context controller predicts mixer parameters, a differentiable DJ mixer blends them, and a discriminator judges the mix.
FIG — controller, mixer, discriminator

About

I'm a PhD student in the Data Science Degree Program at National Taiwan University and Academia Sinica, working with Li Su and Yi-Hsuan Yang. My research is machine learning for music, on both the listening side and the making side.

Lately that has meant a lot of DJ work on pop, where the vocals make everything harder, and before that drum transcription and source separation with latent diffusion.

Before the PhD I worked on electronic dance music classification at Academia Sinica, and as a machine learning engineer at CloudMile, on customer lifetime value models and LLM projects.

After hours

I DJ for fun. I also play a lot of basketball: school teams at NSYSU and NTPU, and a few amateur teams on weekends. Before that it was soccer, on the school team through junior high.