Junhyeok Park

IMG_4836.JPG

Seoul, South Korea

Hi👋, this is Junhyeok! I graduated with a double major in Computer Science and French Language & Literature, and I am currently working at MLAI Lab@Yonsei as a research intern. My research focuses on mechanistic interpretability and AI safety.

Before joining the lab, my work centered on reproductions and independent experiments probing how representational geometry relates to model behavior. I reproduced refusal direction ablation to verify that a single direction in the residual stream mediates refusal, examined linear truth structure across model families and scales using PCA, probing, and causal intervention, and investigated query-conditioned execution heads that causally support task execution in in-context learning. Each of these reads an internal state through a probe or a direction fixed in advance, for a behavior chosen in advance.

Currently, I am interested in methods that translate the internal activations of LLMs into natural language descriptions, such as natural language autoencoders, activation oracles, and LatentQA-style decoding. I am particularly interested in improving the faithfulness of these explanations. If you are interested in these topics, feel free to reach out!

news

Aug 20, 2026 One paper has been accepted to Findings of EMNLP 2026!
May 21, 2026 I’m joining MLAI Lab@Yonsei as a research intern!
Jan 24, 2026 Setting up my personal website!

latest posts

selected publications