Skip to content

Appendix F: References ​

Citations in the body are numbered and link directly to the source below.

#Reference
(1)Russell, S., & Norvig, P. (2020). Artificial Intelligence: A Modern Approach (4th ed.). Pearson.
(2)Feng, K. J. K., McDonald, D. W., & Zhang, A. X. (2025). "Levels of Autonomy for AI Agents." arXiv:2506.12469.
(3)SAE International (2022). SAE J3016: Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles.
(4)Wei, J. et al. (2022). "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models." arXiv:2201.11903.
(5)Ji, Z. et al. (2023). "Survey of Hallucination in Natural Language Generation." arXiv:2202.03629.
(6)Valmeekam, K., Marquez, M., Olmo, A., Sreedharan, S., & Kambhampati, S. (2022). "PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change." arXiv:2206.10498.
(7)Cemri, M., Pan, M. Z., Yang, S., Agrawal, L. A., et al. (2025). "Why Do Multi-Agent LLM Systems Fail?" arXiv:2503.13657.
(8)Tian, K., Mitchell, E., Zhou, A., et al. (2023). "Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback." arXiv:2305.14975.
(9)Sennrich, R., Haddow, B., & Birch, A. (2016). "Neural Machine Translation of Rare Words with Subword Units." arXiv:1508.07909.
(10)Liu, N. F. et al. (2023). "Lost in the Middle: How Language Models Use Long Contexts." arXiv:2307.03172.
(11)Wang, Q., Fu, Y., Cao, Y., et al. (2023). "Recursively Summarizing Enables Long-Term Dialogue Memory in Large Language Models." arXiv:2308.15022.
(12)Lewis, P. et al. (2020). "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks." arXiv:2005.11401.
(13)Willard, B. T., & Louf, R. (2023). "Efficient Guided Generation for Large Language Models." arXiv:2307.09702.
(14)Qin, Y. et al. (2023). "ToolLLM: Facilitating Large Language Models to Master 16000+ Real-World APIs." arXiv:2307.16789.
(15)Patil, S. G., Zhang, T., Wang, X., & Gonzalez, J. E. (2023). "Gorilla: Large Language Model Connected with Massive APIs." arXiv:2305.15334.
(16)Zhang, K. et al. (2025). "LLM Agents Should Employ Security Principles." arXiv:2505.24019.
(17)Zhu, J. et al. (2025). "MiniScope: A Least Privilege Framework for Authorizing Tool Calling Agents." arXiv:2512.11147.
(18)Yang, J., Jimenez, C. E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., & Press, O. (2024). "SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering." arXiv:2405.15793.
(19)Yao, S. et al. (2022). "ReAct: Synergizing Reasoning and Acting in Language Models." arXiv:2210.03629.
(20)Zhang, Z., Bo, X., Ma, C., et al. (2024). "A Survey on the Memory Mechanism of Large Language Model based Agents." arXiv:2404.13501.
(21)Packer, C. et al. (2023). "MemGPT: Towards LLMs as Operating Systems." arXiv:2310.08560.
(22)Vaswani, A. et al. (2017). "Attention Is All You Need." arXiv:1706.03762.
(23)Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., & Fritz, M. (2023). "Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection." arXiv:2302.12173.
(24)Wei, A., Haghtalab, N., & Steinhardt, J. (2023). "Jailbroken: How Does LLM Safety Training Fail?" arXiv:2307.02483.
(25)OpenTelemetry GenAI Semantic Conventions (spec). open-telemetry/semantic-conventions-genai.
(26)Dong, L., Lu, Q., & Zhu, L. (2024). "AgentOps: Enabling Observability of LLM Agents." arXiv:2411.05285.
(27)Zheng, L., Chiang, W.-L., Sheng, Y., et al. (2023). "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena." arXiv:2306.05685.
(28)Jimenez, C. E. et al. (2024). "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" arXiv:2310.06770.
(29)Mialon, G. et al. (2023). "GAIA: a benchmark for General AI Assistants." arXiv:2311.12983.
(30)Liu, X. et al. (2023). "AgentBench: Evaluating LLMs as Agents." arXiv:2308.03688.
(31)Zhou, S. et al. (2023). "WebArena: A Realistic Web Environment for Building Autonomous Agents." arXiv:2307.13854.
(32)Xie, T. et al. (2024). "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments." arXiv:2404.07972.
(33)Hendrycks, D. et al. (2020). "Measuring Massive Multitask Language Understanding." arXiv:2009.03300.
(34)Cobbe, K. et al. (2021). "Training Verifiers to Solve Math Word Problems." arXiv:2110.14168.
(35)Chen, M. et al. (2021). "Evaluating Large Language Models Trained on Code." arXiv:2107.03374.
(36)Rein, D. et al. (2023). "GPQA: A Graduate-Level Google-Proof Q&A Benchmark." arXiv:2311.12022.
(37)Zhou, J. et al. (2024). "Instruction-Following Evaluation for Large Language Models." arXiv:2311.07911.
(38)Suzgun, M. et al. (2022). "Challenging BIG-Bench Tasks and Whether CoT Can Solve Them." arXiv:2210.09261.
(39)Brown, T. B. et al. (2020). "Language Models are Few-Shot Learners." arXiv:2005.14165.
(40)Wang, L. et al. (2023). "Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models." arXiv:2305.04091.
(41)Yao, S. et al. (2023). "Tree of Thoughts: Deliberate Problem Solving with Large Language Models." arXiv:2305.10601.
(42)Madaan, A. et al. (2023). "Self-Refine: Iterative Refinement with Self-Feedback." arXiv:2303.17651.
(43)Shinn, N. et al. (2023). "Reflexion: Language Agents with Verbal Reinforcement Learning." arXiv:2303.11366.
(44)Zou, H. P. et al. (2025). "LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey." arXiv:2505.00753.
(45)Qian, C. et al. (2023). "ChatDev: Communicative Agents for Software Development." arXiv:2307.07924.
(46)Park, J. S. et al. (2023). "Generative Agents: Interactive Simulacra of Human Behavior." arXiv:2304.03442.
(47)Schembri, P.-A. (2026). "Teaching Large Language Models how to generate workflows (2/n)." Advanced Stack.