Personalization & GraphRAG
Connecting users, documents, products, reviews, and conversation history into structured evidence for grounded AI.
Researcher · Engineer · Builder
ML Researcher · NLP Engineer · Lifelong Learner
I build personalized AI systems, with a goal of advancing computational psychiatry and making mental health support more accessible.

I'm an independent NLP researcher and engineer with an M.S. in Natural Language Processing from UC Santa Cruz.
I work across startup and research environments on RAG pipelines, knowledge graphs, personalization, information extraction, and LLM evaluation—especially where accuracy, safety, and context matter.
Connecting users, documents, products, reviews, and conversation history into structured evidence for grounded AI.
Applying NLP to patient-authored text, medication reviews, clinical narratives, and mental health contexts.
Modeling intent, subtext, narrative structure, and interpersonal dynamics beyond surface-level text.
Testing groundedness, interpretability, and failure modes where accuracy, safety, and context matter.
Steven Au et al.
PACLIC 2025 · ACL Anthology
Authors
Steven Au, Cameron Dimacali, Ojasmitha Pedirappagari, Namyong Park, Franck Dernoncourt, Yu Wang, Nikos Kanakaris, Hanieh Deilamsalehy, Ryan A. Rossi, Nesreen K. Ahmed
Abstract
We introduce Personalized Graph-based Retrieval-Augmented Generation (PGraphRAG), a framework that uses user-centric knowledge graphs to improve personalization in cold-start and sparse-data settings. Across diverse tasks, graph-based retrieval improves both relevance and generation quality, with average ROUGE-1 gains of 14.8% on long-text and 4.6% on short-text generation.
Steven Au
NLP4MusA at EACL 2026
Abstract
MIDI-PHOR is a MIDI-first framework that converts symbolic music into structured, queryable representations for reasoning. It distills each piece into symbolic, time-series, and instrument-role graph views, producing evidence-linked claims and reducing hallucinations compared with raw-MIDI baselines.
Neng Wan, Steven Au, Esha Ubale, Decker Krogh
SemEval 2024 · ACL Anthology
Abstract
We describe SemEval-2024 Task 10: EDiReF, consisting of three subtasks involving emotion in conversation across Hinglish code-mixed and English datasets. We deployed a BERT model for emotion recognition and two GRU-based models for emotion flip reasoning, achieving F1 scores of 0.45, 0.79, and 0.68 across the three subtasks.

AI companion for reflection, planning, and personalized recommendations.

Interactive graph of my research, projects, skills, and connections.

NLP pipeline for patient-authored drug reviews and health-related text.
University of California, Santa Cruz
University of California, Santa Cruz
Bytes of Mind Lab, Icahn School of Medicine at Mount Sinai
New York, NY (Remote) | Sep. 2025 – PresentBuilding personalized review-generation benchmarks and healthcare knowledge graphs from patient, medication, adverse-effect, and clinical-trial evidence.
Eternos Inc. / Uare.ai
Remote | Apr. 2025 – Aug. 2025Built multimodal RAG and retrieval pipelines for digital-twin workflows across AWS, Databricks, and long-form conversational data.
Baskin Engineering, UC Santa Cruz
Santa Cruz, CA | Jun. 2024 – Jan. 2025Developed RAG-supported learning experiences and maintained accessible research and engineering web content.
Intel Labs
Santa Clara, CA | May 2024 – Sep. 2024Researched personalized product-review generation using user-item graphs, retrieval baselines, and sparse-user evidence.
I'm open to research collaborations and applied opportunities in computational psychiatry and personalized mental health.