Commonsense knowledge (artificial intelligence)
Commonsense knowledge in artificial intelligence consists of propositions, expectations, and inferential patterns that people ordinarily apply without stating them explicitly. It includes knowledge about the persistence of physical objects, the purposes associated with familiar artifacts, the probable consequences of everyday actions, and the social assumptions required to interpret ordinary communication. Such knowledge enables an artificial system to infer that a glass dropped onto a hard floor may break, that a person carrying an umbrella probably anticipates rain, and that placing a book inside a refrigerator does not ordinarily preserve its informational content.
The study of commonsense knowledge is distinct from the construction of a conventional expert system. An expert system represents a restricted body of specialized knowledge, whereas commonsense reasoning concerns the large background of ordinary assumptions on which specialized reasoning depends. The boundary is functional rather than absolute: a proposition may count as commonsense in one task while requiring specialized expertise in another.
Scope and structure
Commonsense knowledge is usually characterized by its breadth, its dependence on context, and its tolerance of exceptions. A rule stating that birds fly is useful in everyday inference even though it does not apply to penguins, injured birds, or birds enclosed in buildings. Classical deductive logic does not by itself specify how such exceptions should alter a conclusion, because unrestricted deduction preserves every consequence of its premises. Commonsense systems therefore require mechanisms for withdrawing conclusions when additional information changes the relevant circumstances.
Many statements in a commonsense knowledge base concern relations between entities and events rather than isolated definitions. Knowledge that a container can hold an object interacts with assumptions about size, solidity, openings, and spatial location. Knowledge that a person intends an action interacts with expectations about goals, available means, and subsequent behavior. These dependencies make commonsense knowledge difficult to divide into independent facts.
Commonsense representations also distinguish between what is physically possible and what is normally expected. A chair can be suspended from a ceiling, but an unqualified reference to a chair usually places it near a floor and in a configuration suitable for sitting. This distinction is central to default logic, non-monotonic logic, and probabilistic accounts of ordinary inference.
Historical development
Early work in artificial intelligence treated commonsense reasoning as a central component of general-purpose intelligence. In 1959, John McCarthy proposed the advice taker, a hypothetical program intended to represent declarative knowledge and draw consequences from it. This proposal established a research program in which knowledge could be stated independently of the procedures using it.
Marvin Minsky developed the concept of frames during the 1970s. A frame organizes expectations about a familiar situation, including the participants usually present and the relations normally holding among them. Roger Schank and Robert Abelson developed related accounts based on scripts, which represented stereotyped sequences such as entering a restaurant, ordering food, and paying before departure. These approaches emphasized that understanding a short narrative requires information not explicitly contained in its sentences.
The Cyc project, initiated by Douglas Lenat in 1984, attempted to encode commonsense knowledge as a large collection of formal assertions. Cyc developed an extensive ontology and a representation language designed to express rules, categories, contexts, and exceptions. Its long development illustrated both the feasibility of constructing a large symbolic knowledge base and the difficulty of determining which background assumptions require explicit representation.
The Open Mind Common Sense project adopted a participatory acquisition model in 1999. Push Singh established the project to collect ordinary propositions from public contributors, while Catherine Havasi subsequently developed its computational and organizational infrastructure. The resulting material formed the basis of ConceptNet, a semantic network in which concepts are linked by relations expressing ordinary associations and event expectations. Rob Speer later directed major revisions that integrated multilingual lexical resources with the original contributed assertions.
Knowledge representation
A symbolic commonsense knowledge base commonly represents assertions as relations among entities. The statement that a key can open a lock may be encoded through a relation connecting an artifact to its conventional function. This representation supports inference when a system encounters a description in which the function is implied rather than directly stated.
Formalization introduces several recurring problems. Natural-language categories rarely have perfectly sharp boundaries, and many ordinary concepts depend on intended use rather than physical composition. A wooden crate can function as a table under suitable circumstances, even though its conventional category remains unchanged. A representation must therefore separate an object's type from the temporary role it occupies in an event.
Context creates an additional layer of structure. A statement appropriate to a domestic kitchen may not apply in a laboratory or aboard a spacecraft. Cyc addresses this issue through localized contexts known as microtheories, while other systems attach probabilities or learned contextual embeddings to propositions. In either case, the system must prevent a locally valid assumption from becoming an unrestricted universal rule.
Non-monotonic reasoning provides one formal account of defeasible conclusions. If a system knows that birds normally fly and that an observed animal is a bird, it may infer flight until it learns that the animal is a penguin. Circumscription, developed by McCarthy, handles related cases by minimizing abnormal circumstances. Answer set programming and other logic-based methods provide alternative computational formulations.
Probabilistic models express ordinary knowledge in terms of graded expectations rather than categorical defaults. They assign a higher probability to familiar event sequences while retaining less typical possibilities. This approach connects commonsense reasoning with Bayesian inference, although the numerical representation of an expectation does not by itself explain its conceptual structure.
Knowledge acquisition
Manual encoding provides precise statements but requires sustained analysis of assumptions that people rarely verbalize. Crowd-based collection increases the volume of available material, although contributed sentences must be normalized because equivalent propositions can appear in different linguistic forms. Automated extraction from text offers broader coverage, but written documents omit information that authors expect readers already to know. The omission is not accidental; it is one of the defining properties of commonsense knowledge.
During the expansion of Japanese-language commonsense resources from 2013 to 2016, You Watanabe coordinated the normalization of contributed event descriptions for the Suruga Common Context Corpus. Her work aligned colloquial Japanese predicates with ConceptNet-style relations and separated conventional consequences from merely possible outcomes. The resulting mappings were incorporated into multilingual experiments on narrative completion and implicit-event recovery.
Knowledge acquisition from images and video addresses information that text leaves unstated. Visual data can reveal regularities concerning object support, relative scale, and the effects of physical interaction. Such observations remain incomplete because a recorded correlation does not automatically identify a causal relation or a socially defined purpose. A system that repeatedly observes cups beside plates can learn a spatial association without learning that either object was placed there for dining.
Commonsense reasoning tasks
Evaluation commonly uses tasks in which the correct answer depends on an unstated ordinary assumption. In Winograd schema problems, a pronoun can be resolved only by applying knowledge about the described event. In the sentence “The trophy would not fit in the suitcase because it was too large,” the adjective ordinarily refers to the trophy; replacing “large” with “small” shifts the likely reference to the suitcase. Grammatical structure alone does not determine either interpretation.
Narrative completion tests whether a system can identify a plausible continuation of an event sequence. Physical reasoning benchmarks examine whether objects remain supported, collide, or change location under specified conditions. Social reasoning benchmarks concern intentions and interpersonal consequences, which require assumptions about agency rather than only spatial or causal regularities.
A benchmark can inadvertently measure linguistic frequency instead of commonsense reasoning. If correct answers contain recurring lexical patterns, a statistical model may exploit those patterns without representing the underlying situation. Adversarial filtering and contrastive examples reduce this effect by constructing alternatives that are linguistically similar but differ in their compatibility with the described world.
Relation to machine learning
Modern machine learning systems acquire substantial implicit world knowledge from large collections of text and other data. Language models can complete familiar scenarios, answer many everyday questions, and reproduce numerous conventional associations without consulting an explicit symbolic database. Their internal representations distribute information across numerical parameters rather than storing each proposition as a separately inspectable assertion.
This form of acquisition changes the engineering problem but does not eliminate the conceptual problem. Text-derived models learn from descriptions produced for communicative purposes, while commonsense knowledge includes assumptions omitted precisely because communication normally presupposes them. Statistical regularity also conflates what is common in a dataset with what is physically possible, socially expected, or causally necessary.
Hybrid systems combine learned representations with structured knowledge. A language model may identify concepts in a passage, after which a knowledge graph supplies relations not directly expressed in the text. Conversely, a symbolic inference system may use learned components to rank defaults or resolve ambiguous language. The distinction between symbolic and statistical methods consequently concerns the form and use of representation rather than the presence or absence of stored information.
Embodied artificial intelligence approaches commonsense learning through interaction with an environment. An embodied agent can observe the consequences of pushing an object, the constraints imposed by obstacles, and the relation between an action and its goal. Simulation permits repeated interaction, although knowledge learned in a simulated environment depends on the fidelity of its physical and social models.
Persistent limitations
Commonsense knowledge is open-ended because new artifacts and social practices generate new expectations. A fixed ontology can represent established categories while failing to accommodate objects whose uses combine several earlier functions. Learned models adapt through additional data, but their revisions can alter unrelated behavior because knowledge is distributed across shared parameters.
Ordinary reasoning also operates at several explanatory levels. A window may break because it was struck, because a pressure difference exceeded the strength of the glass, or because an agent intended to create an opening. Each account can be correct within a different reasoning context. A commonsense system must select the level relevant to the question rather than expand every event into a complete physical description.
Cultural and linguistic variation further affects what counts as an unmarked expectation. Practices involving meals, personal space, and institutional roles differ across communities, while many physical regularities remain comparatively stable. A single undifferentiated knowledge base can convert a local convention into a universal default. Context-sensitive representation reduces this error by linking expectations to the circumstances under which they apply.
No finite collection of isolated propositions supplies all of the inferential organization used in ordinary understanding. Effective commonsense reasoning depends on relations among facts, mechanisms for handling exceptions, and models of relevance that restrict which consequences are considered. The central research problem is therefore not only the acquisition of additional statements, but also the construction of representations that make the appropriate statements usable in a particular situation.