CV

Contact Information

Name Shreya Nupur Shakya
Professional Title Natural Language Processing & Large Language Models
Email snshakya2025@gmail.com

Professional Summary

Computer Science graduate interested in large language models, efficient and adaptive reasoning, AI agents, multilingual NLP, multimodal learning, and model evaluation.

Experience

  • 2022 - 2025

    Tucson, AZ

    Course Instructor, Department of Computer Science
    University of Arizona
    • Independently led undergraduate courses in Software Development and Web Programming.
    • Delivered lectures and hands-on instruction in programming, software engineering, and web technologies.
    • Designed assignments, assessments, and grading rubrics; evaluated student performance and held office hours.
    • CSC 210 – Software Development: Summer 2025
    • CSC 337 – Web Programming: Summer 2023, Summer 2022
  • 2022 - 2025

    Tucson, AZ

    Graduate Teaching Assistant, Department of Computer Science
    University of Arizona
    • Assisted instructors with course administration and assessment.
    • Set up Gradescope and D2L, organized grading workflows, held office hours, and contributed to grading rubrics when needed.
    • CSC 337 – Web Programming: Fall 2025, Spring 2022
    • CSC 346 – Cloud Computing: Spring 2025
    • CSC 452 – Operating Systems: Fall 2023, Fall 2022
    • CSC 483/583 – Text Retrieval and Web Search: Spring 2023
  • 2021 - 2021

    Tucson, AZ

    Graduate Research Assistant
    University of Arizona
    Advisor: Dr. Katherine E. Isaacs
    • Explored visualization methods for understanding compiler optimization behavior through static and dynamic program analysis.
    • Developed an interactive visualization system using Dyninst APIs and D3.js to expose code-optimization behavior and support analysis of program transformations.
  • Tucson, AZ

    Research Contributor – Interpreting Indirect Answers to Yes-No Questions in Multiple Languages
    University of Arizona
    Advisor: Dr. Eduardo Blanco
    • Contributed to research on multilingual pragmatic language understanding, focusing on how indirect answers to yes-no questions are interpreted across languages.
    • Created and validated Nepali language data using native-language judgments of context-dependent answers.
    • Supported evaluation of cross-lingual transfer across eight languages.
    • Co-authored the resulting work published in Findings of EMNLP 2023.
  • Tucson, AZ

    Research Contributor – SexTok
    University of Arizona
    Advisor: Dr. Mihai Surdeanu
    • Supported development of a 1,000-video multimodal dataset for distinguishing sex-educational, sexually suggestive, and other TikTok content.
    • Annotated context-sensitive social-media content using visual and linguistic cues to support multimodal content classification.
    • Recognized in the acknowledgements of the resulting Findings of ACL 2023 paper for data annotation contributions.

Interests

Research Interests: LLM Reasoning and Evaluation, AI Agents, Efficient and Adaptive Reasoning, Multimodal Learning, Multilingual NLP

Education

  • 2021 - 2025

    Tucson, AZ

    Master of Science
    University of Arizona
    Computer Science
    • Graduate Certificate in Natural Language Processing
    • Medical leave [Spring 2024, Fall 2024]
  • 2017 - 2021

    Bangalore, India

    Bachelor of Technology
    Jain University
    Computer Science and Engineering
    • Rank: 2/~500

Publications

Projects

  • Adaptive Reasoning and Tool Routing from LLM Hidden States

    Predicting inference-time resource needs from pre-generation representations.

    • Studied whether hidden states can predict when additional reasoning or external tool access is needed before generation.
    • Used Qwen3-1.7B and layer-wise linear probes to predict model-adaptive reasoning and tool necessity.
    • Achieved peak AUROC of 0.881 for reasoning necessity and 0.963 for tool necessity.
    • Evaluated selective routing policies for allocating Thinking and tool access under fixed resource budgets.
  • Minimal Chain-of-Thought Reasoning in Low-Resource Languages

    LING 582 – Advanced Statistical Natural Language Processing, University of Arizona

    • Studied whether shorter reasoning traces can preserve LLM reasoning performance in a low-resource language while reducing inference cost.
    • Compared No-CoT, Minimal-CoT, and Standard-CoT using Llama-3-8B-Instruct and Qwen2.5-7B-Instruct across English and Nepali mathematical reasoning.
    • Found that Minimal-CoT retained approximately 95–100% of Standard-CoT accuracy in three of four model-language settings while using 68–81% of its token budget.
    • Analyzed failures involving truncation, semantic grounding, reasoning drift, and repetition.
  • Further Pre-training RoBERTa for Negation in Neural Information Retrieval

    CSC 583 – Text Retrieval and Web Search, University of Arizona

    • Examined whether negation-specific further pre-training improves neural retrievers’ sensitivity to semantic negation.
    • Evaluated RoBERTa-large on NevIR under no-supervision and fine-tuning settings.
    • Compared negation-focused pre-training methods including NSP, NSPP, and commonsense-based objectives.
    • Achieved a NevIR score of 80.8 with NSP-pretrained RoBERTa-large after fine-tuning, outperforming the reported RankGPT o3-mini result by 3.5 points.
  • Interpreting Indirect Answers to Yes-No Questions in Multiple Languages

    Multilingual pragmatic language understanding research, University of Arizona

    • Studied how indirect answers to yes-no questions are interpreted across languages.
    • Contributed to multilingual data collection, validation, and curation, including the Nepali benchmark.
    • Supported evaluation of cross-lingual transfer across eight languages.
    • Resulting work was published in Findings of EMNLP 2023.
  • SexTok – Separating Sex Education from Suggestive Content on TikTok

    Multimodal social-media content moderation research, University of Arizona

    • Contributed to dataset annotation for a 1,000-video multimodal TikTok dataset.
    • Helped distinguish sex-educational, sexually suggestive, and other content using visual and linguistic context.
    • Resulting research was published in Findings of ACL 2023.

Skills

Programming (): Python, Java, C/C++, SQL, JavaScript
ML/NLP Libraries (): PyTorch, TensorFlow, Hugging Face Transformers, scikit-learn
Tools and Platforms (): Git, Docker, Linux, Jupyter
Databases (): PostgreSQL, MySQL, MongoDB