Appearance
Appendix F: References
Citations in the body are numbered and link directly to the source below.
| # | Reference |
|---|---|
| (1) | Russell, S., & Norvig, P. (2020). Artificial Intelligence: A Modern Approach (4th ed.). Pearson. |
| (2) | Feng, K. J. K., McDonald, D. W., & Zhang, A. X. (2025). "Levels of Autonomy for AI Agents." arXiv:2506.12469. |
| (3) | SAE International (2022). SAE J3016: Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles. |
| (4) | Wei, J. et al. (2022). "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models." arXiv:2201.11903. |
| (5) | Ji, Z. et al. (2023). "Survey of Hallucination in Natural Language Generation." arXiv:2202.03629. |
| (6) | Valmeekam, K., Marquez, M., Olmo, A., Sreedharan, S., & Kambhampati, S. (2022). "PlanBench: An Extensible Benchmark for Evaluating Large Language Models on Planning and Reasoning about Change." arXiv:2206.10498. |
| (7) | Cemri, M., Pan, M. Z., Yang, S., Agrawal, L. A., et al. (2025). "Why Do Multi-Agent LLM Systems Fail?" arXiv:2503.13657. |
| (8) | Tian, K., Mitchell, E., Zhou, A., et al. (2023). "Just Ask for Calibration: Strategies for Eliciting Calibrated Confidence Scores from Language Models Fine-Tuned with Human Feedback." arXiv:2305.14975. |
| (9) | Sennrich, R., Haddow, B., & Birch, A. (2016). "Neural Machine Translation of Rare Words with Subword Units." arXiv:1508.07909. |
| (10) | Liu, N. F. et al. (2023). "Lost in the Middle: How Language Models Use Long Contexts." arXiv:2307.03172. |
| (11) | Wang, Q., Fu, Y., Cao, Y., et al. (2023). "Recursively Summarizing Enables Long-Term Dialogue Memory in Large Language Models." arXiv:2308.15022. |
| (12) | Lewis, P. et al. (2020). "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks." arXiv:2005.11401. |
| (13) | Willard, B. T., & Louf, R. (2023). "Efficient Guided Generation for Large Language Models." arXiv:2307.09702. |
| (14) | Qin, Y. et al. (2023). "ToolLLM: Facilitating Large Language Models to Master 16000+ Real-World APIs." arXiv:2307.16789. |
| (15) | Patil, S. G., Zhang, T., Wang, X., & Gonzalez, J. E. (2023). "Gorilla: Large Language Model Connected with Massive APIs." arXiv:2305.15334. |
| (16) | Zhang, K. et al. (2025). "LLM Agents Should Employ Security Principles." arXiv:2505.24019. |
| (17) | Zhu, J. et al. (2025). "MiniScope: A Least Privilege Framework for Authorizing Tool Calling Agents." arXiv:2512.11147. |
| (18) | Yang, J., Jimenez, C. E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., & Press, O. (2024). "SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering." arXiv:2405.15793. |
| (19) | Yao, S. et al. (2022). "ReAct: Synergizing Reasoning and Acting in Language Models." arXiv:2210.03629. |
| (20) | Zhang, Z., Bo, X., Ma, C., et al. (2024). "A Survey on the Memory Mechanism of Large Language Model based Agents." arXiv:2404.13501. |
| (21) | Packer, C. et al. (2023). "MemGPT: Towards LLMs as Operating Systems." arXiv:2310.08560. |
| (22) | Vaswani, A. et al. (2017). "Attention Is All You Need." arXiv:1706.03762. |
| (23) | Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., & Fritz, M. (2023). "Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection." arXiv:2302.12173. |
| (24) | Wei, A., Haghtalab, N., & Steinhardt, J. (2023). "Jailbroken: How Does LLM Safety Training Fail?" arXiv:2307.02483. |
| (25) | OpenTelemetry GenAI Semantic Conventions (spec). open-telemetry/semantic-conventions-genai. |
| (26) | Dong, L., Lu, Q., & Zhu, L. (2024). "AgentOps: Enabling Observability of LLM Agents." arXiv:2411.05285. |
| (27) | Zheng, L., Chiang, W.-L., Sheng, Y., et al. (2023). "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena." arXiv:2306.05685. |
| (28) | Jimenez, C. E. et al. (2024). "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" arXiv:2310.06770. |
| (29) | Mialon, G. et al. (2023). "GAIA: a benchmark for General AI Assistants." arXiv:2311.12983. |
| (30) | Liu, X. et al. (2023). "AgentBench: Evaluating LLMs as Agents." arXiv:2308.03688. |
| (31) | Zhou, S. et al. (2023). "WebArena: A Realistic Web Environment for Building Autonomous Agents." arXiv:2307.13854. |
| (32) | Xie, T. et al. (2024). "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments." arXiv:2404.07972. |
| (33) | Hendrycks, D. et al. (2020). "Measuring Massive Multitask Language Understanding." arXiv:2009.03300. |
| (34) | Cobbe, K. et al. (2021). "Training Verifiers to Solve Math Word Problems." arXiv:2110.14168. |
| (35) | Chen, M. et al. (2021). "Evaluating Large Language Models Trained on Code." arXiv:2107.03374. |
| (36) | Rein, D. et al. (2023). "GPQA: A Graduate-Level Google-Proof Q&A Benchmark." arXiv:2311.12022. |
| (37) | Zhou, J. et al. (2024). "Instruction-Following Evaluation for Large Language Models." arXiv:2311.07911. |
| (38) | Suzgun, M. et al. (2022). "Challenging BIG-Bench Tasks and Whether CoT Can Solve Them." arXiv:2210.09261. |
| (39) | Brown, T. B. et al. (2020). "Language Models are Few-Shot Learners." arXiv:2005.14165. |
| (40) | Wang, L. et al. (2023). "Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models." arXiv:2305.04091. |
| (41) | Yao, S. et al. (2023). "Tree of Thoughts: Deliberate Problem Solving with Large Language Models." arXiv:2305.10601. |
| (42) | Madaan, A. et al. (2023). "Self-Refine: Iterative Refinement with Self-Feedback." arXiv:2303.17651. |
| (43) | Shinn, N. et al. (2023). "Reflexion: Language Agents with Verbal Reinforcement Learning." arXiv:2303.11366. |
| (44) | Zou, H. P. et al. (2025). "LLM-Based Human-Agent Collaboration and Interaction Systems: A Survey." arXiv:2505.00753. |
| (45) | Qian, C. et al. (2023). "ChatDev: Communicative Agents for Software Development." arXiv:2307.07924. |
| (46) | Park, J. S. et al. (2023). "Generative Agents: Interactive Simulacra of Human Behavior." arXiv:2304.03442. |
| (47) | Schembri, P.-A. (2026). "Teaching Large Language Models how to generate workflows (2/n)." Advanced Stack. |