Data science · NLP · Applied AI

I build data and AI systems that turn complex information into useful decisions.

I’m Stella (Juanru) Zhang, a Cornell Information Science graduate and incoming NYU Data Science master’s student. My work spans human–LLM value alignment, natural-language interfaces for structured data, machine learning, analytics, and software engineering.

Cornell University B.A. Information Science, 2026
New York University Incoming M.S. Data Science, Fall 2026
Focus NLP, AI systems, data analysis

Selected work

Projects

Experience & research

Experience

May 2025–Present

Ithaca, NY

Research Assistant

Cornell Human–LLM Value Alignment Research

  • Designed an experimental framework to study how LLM-generated constructive comments affect the trajectory and rhetoric of online discussions.
  • Conducted large-scale R and SQL analysis of value alignment, cultural homogenization, and model-generated interventions.
  • Co-authored the ACM CHI 2026 paper “LLMs Homogenize Values in Constructive Arguments on Value-Laden Topics” , examining how LLM rewrites can suppress some values while amplifying others in contentious online discussions.

Jun–Oct 2024

Remote, US

Software Engineer

Convoloo — NLP Internship

  • Built a LangChain and Gemini question-answering system that translated natural language into SQL for structured patient-health datasets.
  • Achieved 95% accuracy on basic information-retrieval queries and integrated MySQL and SQLite with AI inference pipelines.
  • Developed conversational retrieval chains and output templates that made generated queries and results easier to inspect and debug.

Jun–Aug 2024

Beijing, China

Data Analyst

Xiaomi — Data Analysis Internship

  • Validated weekly product data for more than 5,000 SKUs across labels, costs, revenue, and inventory status.
  • Built four Power BI dashboards tracking product lifecycles, regional sales, and inventory turnover.
  • Used SQL and Excel to improve reporting workflows across six regions, reducing report-construction time by 30%.
Earlier experience

Software Engineer / Data Analyst

Hisense Transport Statistics Management

Developed debugging protocols, integrated Alibaba Sentinel for request regulation, and supported a cross-functional peer-review process.

Programmer / Data Analyst

Baidu Data Analysis Mentorship

Evaluated XGBoost, Random Forest, Logistic Regression, and Naïve Bayes models on more than 10,000 website visits, achieving over 95% classification accuracy.

Technical toolkit

Skills

Languages

Python, SQL, R, Java, JavaScript, C++, OCaml, HTML, CSS

Data & machine learning

cikit-learn, XGBoost, random forest, gradient boosting, logistic regression, k-means, PCA, cross-validation, hyperparameter tuning, feature engineering, MySQL, SQLite, LangChain, retrieval-augmented generation (RAG), text-to-SQL, prompt engineering, large language models (LLMs)

Analytics and Visualization

pandas, NumPy, Power BI, Tableau, Matplotlib, seaborn, ggplot2, Microsoft Excel, A/B testing, statistical analysis

Developer Tools

GitHub, Jupyter, VS Code, Node.js

Contact

Let’s talk about data, NLP, or applied AI.

I’m interested in research and engineering work where careful analysis and practical software meet.