Sources
References
The intellectual foundations for the concepts developed in this paper.
9 min read
The references below provide the principal intellectual foundations for the concepts developed in this paper. They are organised by area for readability in this version of the white paper. A publication version may instead present them as a single alphabetical bibliography.
S O F T WA R E A R C H I T E C T U R E , M O D U L A R I T Y, A N D C O M P O S I T I O N
Bass, L., Clements, P., & Kazman, R. (2021). Software Architecture in Practice (4th ed.).
Addison-Wesley Professional.
Fielding, R. T. (2000). Architectural Styles and the Design of Network-Based Software
Architectures [Doctoral dissertation, University of California, Irvine].
Gamma, E., Helm, R., Johnson, R., & Vlissides, J. (1994). Design Patterns: Elements of
Reusable Object-Oriented Software. Addison-Wesley.
Lewis, J., & Fowler, M. (2014). Microservices: A definition of this new architectural
term. MartinFowler.com.
Parnas, D. L. (1972). On the criteria to be used in decomposing systems into modules.
Communications of the ACM, 15(12), 1053–1058. DOI: 10.1145/361598.361623.
These works provide the architectural lineage for several ideas used throughout this paper, particularly responsibility-based decomposition, information hiding, encapsulation, stable interfaces, independent capability boundaries, and the organisation of systems from composable parts. The Cognitive Unit is not presented as equivalent to a traditional function, module, object, or service, but it extends a long-standing architectural principle: meaningful responsibility becomes easier to build, reuse, change, and reason about when it has an explicit boundary.
F O U N D AT I O N M O D E L S , L A N G U A G E M O D E L S , T O O L U S E , A N D A G E N T S
Bommasani, R., Hudson, D. A., Adeli, E., Altman, R., Arora, S., von Arx, S., Bernstein, M.
S., Bohg, J., Bosselut, A., Brunskill, E., et al. (2021). On the opportunities and risks of foundation models. arXiv preprint arXiv:2108.07258.
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan,
A., Shyam, P., Sastry, G., Askell, A., et al. (2020). Language models are few-shot learners. Advances in Neural Information Processing Systems, 33, 1877–1901.
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda,
N., & Scialom, T. (2023). Toolformer: Language models can teach themselves to use tools. Advances in Neural Information Processing Systems, 36.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., &
Polosukhin, I. (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30.
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q. V., & Zhou,
D. (2022). Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35, 24824–24837.
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct:
Synergizing reasoning and acting in language models. International Conference on Learning Representations.
These works establish the technological context from which Thoughtware emerges: general-purpose language models, foundation models that transfer across tasks, language-mediated reasoning, interaction with external tools, and systems capable of interleaving reasoning with action. Thoughtware builds above these mechanisms rather than proposing an alternative to them. Its concern is the architectural organisation of the cognitive capabilities that these technologies make possible.
J U D G M E N T, W O R K I N G M E M O R Y, A N D E X P E R T I S E
Baddeley, A. (2000). The episodic buffer: A new component of working memory?
Trends in Cognitive Sciences, 4(11), 417–423. DOI: 10.1016/S1364-6613(00)01538-2.
Ericsson, K. A., Krampe, R. T., & Tesch-Römer, C. (1993). The role of deliberate practice
in the acquisition of expert performance. Psychological Review, 100(3), 363–406. DOI: 10.1037/0033-295X.100.3.363.
Evans, J. St. B. T. (2008). Dual-processing accounts of reasoning, judgment, and social
cognition. Annual Review of Psychology, 59, 255–278. DOI: 10.1146/annurev.psych.59.103006.093629.
Kahneman, D. (2011). Thinking, Fast and Slow. Farrar, Straus and Giroux.
Tversky, A., & Kahneman, D. (1974). Judgment under uncertainty: Heuristics and
biases. Science, 185(4157), 1124–1131. DOI: 10.1126/science.185.4157.1124.
These references inform the paper’s treatment of judgment, limited working context, experience, expertise, and faster versus more deliberative forms of decision-making. The analogy is intentionally limited. Thoughtware does not claim that computational systems reproduce human cognitive architecture, nor that concepts such as System 1 and System 2 map directly onto particular model behaviours. Cognitive science is used here as a conceptual mirror for separating functions that are otherwise easily collapsed into the broad label of intelligence.
H U M A N - A I I N T E R A C T I O N , A U T O M AT I O N , A N D H U M A N O V E R S I G H T
Amershi, S., Weld, D., Vorvoreanu, M., Fourney, A., Nushi, B., Collisson, P., Suh, J., Iqbal,
S., Bennett, P. N., Inkpen, K., Teevan, J., Kikin-Gil, R., & Horvitz, E. (2019). Guidelines for human-AI interaction. Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, 1–13. DOI: 10.1145/3290605.3300233.
Horvitz, E. (1999). Principles of mixed-initiative user interfaces. Proceedings of the
SIGCHI Conference on Human Factors in Computing Systems, 159–166. DOI: 10.1145/302979.303030.
Lee, J. D., & See, K. A. (2004). Trust in automation: Designing for appropriate reliance.
Human Factors, 46(1), 50–80. DOI: 10.1518/hfes.46.1.50_30392.
Parasuraman, R., Sheridan, T. B., & Wickens, C. D. (2000). A model for types and levels
of human interaction with automation. IEEE Transactions on Systems, Man, and Cybernetics — Part A: Systems and Humans, 30(3), 286–297. DOI: 10.1109/3468.844354.
This body of work provides important foundations for the paper’s distinction between capability and authority. Questions about when systems should act autonomously, when humans should intervene, how automation alters human work, how trust should be calibrated, and how mixed-initiative systems divide responsibility long predate foundation models. Thoughtware extends these questions into software architectures in which machine judgment can itself be decomposed, reused, and orchestrated.
E VA L U AT I O N , C A L I B R AT I O N , A N D R E L I A B I L I T Y
Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural
networks. Proceedings of the 34th International Conference on Machine Learning, 1321–1330.
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y.,
Narayanan, D., Wu, Y., Kumar, A., et al. (2022). Holistic evaluation of language models. arXiv preprint arXiv:2211.09110.
Liu, Y., Iter, D., Xu, Y., Wang, S., Xu, R., & Zhu, C. (2023). G-Eval: NLG evaluation using
GPT-4 with better human alignment. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing.
National Institute of Standards and Technology. (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0). NIST AI 100-1.
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D.,
Xing, E. P., Zhang, H., Gonzalez, J. E., & Stoica, I. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. Advances in Neural Information Processing Systems, 36.
These works support the paper’s argument that cognitive systems require forms of assurance beyond deterministic output comparison. They address multidimensional model evaluation, confidence calibration, automated evaluation using language models, agreement with human preferences, evaluator bias, and lifecycle governance. The paper’s concept of Judge CUs builds on this emerging body of work while placing evaluation at the level of explicit cognitive responsibility rather than treating it only as model benchmarking.
P O S I T I O N O F T H I S PA P E R
The concepts Thoughtware, Cognitive Unit, Judgment Terrain, Instruction Surface, Intent Compilation, Submergence Principle, and re-naturalisation, together with the specific architectural relationships proposed among them, are introduced or developed in this paper. The preceding literature provides important foundations and adjacent ideas, but these terms should not be read as established terminology from the cited sources.
The purpose of the references is therefore not to suggest that Thoughtware follows directly from a single existing research tradition. The proposal sits at the intersection
of several: software architecture provides the principles of decomposition and encapsulation; foundation models provide programmable cognitive capability; cognitive science provides useful distinctions around judgment and expertise; human-computer interaction provides a history of allocating responsibility between people and machines; and contemporary evaluation research provides methods for measuring behaviour where deterministic verification is insufficient.
Thoughtware attempts to bring these strands together at the level of software architecture.
Thank You