Fruggia.com
Cover art for The Machine That Guesses

The Machine That Guesses

How Large Language Models Actually Work

  • 8 chapters
  • 56m
  • Artificial Intelligence
  • Free · no sign-up
A language model processes text by breaking it into tokens, then predicts the next word based on patterns it learned during training. This audiobook explains how these systems actually work, covering the algorithms behind tokenization, attention mechanisms, and fine-tuning processes.

The book explores why models sometimes generate false information, known as hallucinations, and how bias appears in artificial intelligence systems. It also discusses the Turing test, creativity in machines, and various benchmarks used to measure language model performance.

Whether you're curious about how AI systems guess what comes next in text or want to understand the real limits of current technology, this straightforward guide offers clear explanations without technical jargon.

Listen

  1. 01 Large language model 7m Download (3.3 MB)
    Read this chapter

    Overview

    A large language model, or LLM, is an artificial intelligence system, usually built as a neural network, trained on enormous amounts of text to handle tasks like understanding and creating human language. These models can write, summarize, translate, and even analyze information across many topics. They power today’s popular chatbots such as ChatGPT, Claude, Gemini, Grok, and DeepSeek. Most LLMs use a structure called the transformer architecture, with GPTs being a common type that learns to predict the next word in a sequence. After being pre-trained this way, they’re often adjusted to better follow instructions and act as helpful assistants. But because their training depends on existing text, if that data contains biases or errors, the results can be unreliable. Researchers try to assess how well these models think, recall facts, and behave safely using benchmark tests.

    History

    Before transformer models emerged in 2017, some language systems were already considered large given the data and computing limits of their time. In the early 1990s, IBM’s statistical models advanced word alignment for translation. By 2001, a smoothed n-gram model trained on 300 million words reached state-of-the-art results. During the 2000s, researchers used web-based text collections to train better statistical models. Then, in 2000, neural networks began replacing n-gram methods. After 2012, deep learning techniques were applied to language tasks, leading to innovations like Word2Vec and LSTM-based seq2seq models. Google shifted to neural machine translation in 2016, using LSTMs. At NeurIPS 2017, Google introduced the transformer architecture in “Attention Is All You Need.” BERT followed in 2018 and quickly became widespread. Although GPT-1 debuted in 2018, GPT-2 in 2019 gained attention for its power. GPT-3 arrived in 2020 and is now only available through API. ChatGPT, released late 2022, sparked public fascination by 2023. GPT-4 in 2023 was praised for accuracy and multimodal abilities. OpenAI did not disclose GPT-4’s structure or parameters. The release of ChatGPT spurred interest across many computer science fields. In 2024, OpenAI introduced the reasoning model o1. Since 2022, open-source models have grown in popularity, starting with BLOOM and LLaMA, though both had usage restrictions. Mistral’s models later offered more permissive licenses. In January 2025, DeepSeek released R1, a 671-billion-parameter model that rivals o1 but costs less. Since 2023, many large language models have been trained to process multiple data types like images or audio. Open-weight models have become increasingly influential since 2023, with community contributions improving their performance via platforms such as Hugging Face.

    Tokenization

    Before any machine learning algorithm can work with text, that text has to be turned into numbers. The process starts by creating a vocabulary, then assigning unique integer indices to each entry in that vocabulary. An embedding is linked to each index. Algorithms like byte-pair encoding and WordPiece handle this conversion. Special tokens act as control characters—like [MASK] in BERT or [UNK] for unknown words. Symbols such as "Ġ" mark whitespace in RoBERTa and GPT, while "##" shows word continuation in BERT. For example, the legacy GPT-3 used a BPE tokenizer to turn text into numerical tokens. Tokenization also helps compress data, since input arrays must be uniform in length, so shorter texts get padded to match the longest one.

    Byte-pair encoding

    Consider a tokenizer using byte-pair encoding. First, every unique character—spaces, punctuation, everything—is treated as a starting set of unigrams. Then, the most commonly occurring pair of adjacent characters gets merged into a bigram, and every instance of that pair is replaced. This process repeats with the newly formed n-grams, merging the most frequent adjacent pairs again, until the vocabulary reaches its target size. Once trained, the tokenizer can break down any text, as long as it only uses characters from the original set.

    Dataset cleaning

    When training large language models, the data used is cleaned by removing low-quality, duplicated, or toxic content. This process helps improve how efficiently the model learns and how well it performs later on. Once a model is trained, it can even be used to clean data for training other models. As more content on the web is being generated by these same language models, future cleaning efforts may need to filter out that kind of material. The challenge is that this artificial content can resemble human-written text, making it hard to spot, yet it's often of lower quality, which hurts the performance of models trained on it.

    Cost

    Training the biggest language models requires serious infrastructure, and the trend has been toward ever-larger systems. For example, in 2019, it cost $50,000 to train GPT-2, a model with 1.5 billion parameters. By 2022, training PaLM, which has 540 billion parameters, set the bill at $8 million. In the same year, Megatron-Turing NLG 530B cost around $11 million to train. The term “large” in “large language model” doesn’t have a fixed meaning, since there’s no official number of parameters that defines what counts as large.

    Fine-tuning

    Most large language models start out as simple next-token predictors, but then they get shaped through a process called fine-tuning. One way this happens is through instruction fine-tuning, which teaches the model to follow directions—like in OpenAI’s InstructGPT from 2022, a version of GPT-3 trained this way. Another method uses reinforcement learning from human feedback, or RLHF. With RLHF, a reward model is trained first to guess what humans prefer in text. Then the language model itself is fine-tuned using that reward model through reinforcement learning to better match those preferences.

    Inference

    When a large language model is used, it’s called inference—basically, you give it some input, like a question or a prompt, and it gives you back output. That output can be text, images, code, video, or even files like CSV, JSON, XML, PDFs, or Word documents. These models are often accessed through chatbot websites or apps, which might be free or require a subscription. You can also interact with them using command-line tools, or by adding them into coding environments like Visual Studio Code. Some setups let you use LLMs directly from your editor, either through built-in tools, external APIs, or by running the model locally on your own device.

  2. 02 Language model benchmark 4m Download (1.7 MB)
    Read this chapter

    Overview

    A language model benchmark is a standard test used to measure how well different language models perform on tasks like understanding and generating text. These tests help compare models based on things such as reasoning, language comprehension, and text creation. Each benchmark includes a dataset with examples and annotations, along with specific metrics to assess performance. The metrics might look at accuracy, but also consider factors like speed, energy use, fairness, trustworthiness, and environmental impact. Such benchmarks are created and kept up by universities, research groups, and companies working in the field to monitor progress over time.

    Types

    Benchmarks for evaluating language models include classical ones like Penn Treebank and BLEU scores, which preceded deep learning, and question-answering tasks ranging from open-book to closed-book QA that became common after GPT-2. Omnibus benchmarks combine multiple tests, while reasoning tasks challenge QA capabilities. Multimodal benchmarks process images or sound with text, such as OCR. Agency benchmarks measure web browsing or image editing abilities. Adversarial benchmarks are designed to trip up models and evolve with performance improvements. Public benchmarks allow open access while private ones prevent training on their data to avoid cheating. Key distinction: datasets have training/validation splits while benchmarks test multiple models without prior training, often lacking corresponding training sets. Some benchmarks like English Gigaword or One Billion Word Benchmark serve as training data themselves, especially during pretraining of modern language models.

    Evaluation

    Benchmarks used to test language models are usually fully automated, which means they can only ask certain kinds of questions. For example, math problems that require proving something are hard to check automatically, but ones with a single integer answer are easy. Programming tasks can be tested by running unit tests with time limits. The results are measured in different ways: for multiple choice or cloze tests, you might see accuracy, precision, recall, or F1 scores. There are also special scores like pass@n, where the model gets n tries per problem and earns a point if any attempt is right. k@n works similarly but only submits k of the n attempts. cons@n gives credit if the most common answer is correct. And for open-ended tasks, metrics like BLEU, ROUGE, METEOR, and others are used to score how well the output matches expected answers.

    Omnibus

    Some language model benchmarks are omnibus, meaning they pull together multiple earlier tests into one larger evaluation. Big-Bench, also known as Beyond the Imitation Game, includes 204 tasks, with a subset called BBH, or Big-Bench Hard. An even tougher version, BBEH, or Big-Bench Extra Hard, replaces those tasks with adversarial versions. GLUE, short for General Language Understanding Evaluation, combines nine smaller tests with over a million items. SuperGLUE updated that set in 2019 to challenge state-of-the-art models, adding eight new tasks like coreference resolution. HELM is a continuously evolving framework from Stanford’s Center for Research on Foundation Models. MMLU measures multitask understanding across 57 subjects with 16,000 questions, later upgraded to MMLU-Pro. CMMLU is the Chinese version, with 1,528 questions covering 67 subjects including some specific to China. MMMLU translates the original MMLU test into 14 languages using human translators.

  3. 03 Algorithmic bias 7m Download (3.3 MB)
    Read this chapter

    Overview

    Algorithmic bias refers to computer systems' consistent tendency to generate unfair results, often favoring certain groups over others contrary to intended function. This bias arises from design decisions or data gathering, labeling, or usage during training, appearing in search engines and social media where it reinforces discrimination based on race, gender, sexuality, or ethnicity. The issue has only recently begun to be tackled in law, with efforts like the European Union's data protection rules taking effect in 2018 and a proposed artificial intelligence act passed in 2024. As these systems grow more influential in shaping society, researchers are increasingly worried about unintended consequences or manipulated data affecting real-world outcomes. Bias may come from cultural assumptions, technical constraints, or unanticipated uses of technology. For example, facial recognition tools have been found to misidentify darker-skinned faces, leading to wrongful arrests. Because these systems are often secretive and complex, it's hard to fully understand or detect their biases. A 2021 survey identified several types of bias—historical, representation, and measurement—that can all lead to unfair results.

    Definitions

    Algorithms are sets of instructions that tell computers how to gather, sort, and analyze data to produce results. As computing power has grown, so have our abilities to store and process massive amounts of information, making advanced technologies like machine learning and AI possible. These systems power search engines, social media, online shopping, and advertising. Scholars are especially interested in how algorithms shape society because they can reflect or reinforce unfair practices. Algorithmic bias refers to consistent mistakes that lead to unequal treatment, such as favoring one group over another. For example, a credit algorithm might reject a loan application without being obviously unfair if it's based on solid financial data. But if it treats nearly identical applicants differently simply because of unrelated factors, and does so repeatedly, then it's biased. This bias can be deliberate or accidental—sometimes it comes from the data used to train the system, which may reflect past human decisions.

    Methods

    Algorithms pick up bias through how data is gathered and organized. When datasets are assembled, human choices shape everything from information collection to inclusion decisions. Programmers determine how to sort and weigh data, allowing their perspectives to influence results. Some systems build data based on human-selected rules, carrying designer biases forward. Others reinforce stereotypes by showing users content similar to others like them. Recommendation engines may rely on faulty links between traits like race or gender and behavior. Simple decisions about what to show or hide can create unexpected outcomes—like a flight-search tool excluding routes not aligned with airline paths. Because algorithms are more certain with more data, they often favor results matching larger groups, pushing smaller or underrepresented voices further into the background.

    Early critiques

    Programs are built from rules set by humans, carrying assumptions and biases of creators, as Joseph Weizenbaum noted in Computer Power and Human Reason. He warned that bias enters through both data and system design, with computers following instructions consistently embodying creators' logic and expectations. Data reflects human choices in what gets selected and included. Weizenbaum compared trusting such systems without understanding them to a tourist finding a hotel room by flipping a coin—success doesn't mean accuracy. In the early 1980s, St. George's Hospital Medical School used a computer system that rejected up to sixty women and minorities annually because their names sounded foreign, following historical patterns of exclusion. As algorithms now rely more on real-world data like facial recognition systems that misidentify darker-skinned women at rates as high as 35%, these biases become harder to ignore. Critics like Cathy O'Neil have shown how automated decisions in policing and credit can reinforce unfair practices while seeming neutral or scientific.

    Contemporary critiques and responses

    Algorithms are often seen as fairer than human choices, but bias still creeps in, and it's hard to predict or analyze. As systems grow more complex, individual designer decisions get lost in layers of code, possibly shaping new patterns over time. Clay Shirky calls this "algorithmic authority," where outputs feel neutral even when they're not. This sense of neutrality can mislead people—like when "trending" news is influenced by more than just popularity. Critics say relying on algorithms shifts responsibility from humans and reduces flexibility. In response, groups like Google and Microsoft have formed working groups such as Fairness, Accountability, and Transparency in Machine Learning. The field has grown into its own area of study with an annual FAccT conference. Some doubt these initiatives can act as true watchdogs if many are funded by the companies they oversee.

    Pre-existing

    Algorithms don't begin with a blank slate; they carry forward societal and institutional biases. When training data reflects existing ideas—deliberately or accidentally—the results mirror those perspectives. The British Nationality Act Program from 1981 automated citizenship decisions based on legal assumptions, encoding those views into its logic even after the law was repealed. Another bias occurs when flawed measures train AI systems. One widely used algorithm predicted health care costs as a stand-in for actual needs, excluding Black patients since they often had lower costs despite being just as ill as White patients. By shifting focus from cost to health needs, researchers nearly doubled the number of Black patients selected for care programs.

    Language bias

    Language bias happens when large language models, trained mostly on English data, end up reflecting an Anglo-American viewpoint as if it were universal. Luo et al.'s research shows this leads to a systematic skew where non-English perspectives are often ignored or treated as less valid. For example, when asked about liberalism, these models tend to focus on human rights and equality—views rooted in Western tradition—while leaving out other interpretations, like those found in Vietnamese or Chinese political thought. The models may also show bias against certain dialects within a language group, further narrowing the range of ideas they present as accurate or important.

    Selection bias

    Selection bias is a problem in how large language models work, where they tend to prefer certain answer options just because of how they're labeled, not because of what they say. This happens because the model has a built-in tendency to pick specific tokens—like "A"—more often than others when generating answers. So if you change the order of the choices, putting the right answer in a different spot, the model’s performance can jump around a lot. That makes it unreliable when used in tests or situations where multiple-choice questions are involved.

  4. 04 Glossary of artificial intelligence 11m Download (4.9 MB)
    Read this chapter

    Overview

    This glossary offers definitions for terms and ideas central to artificial intelligence, including its smaller fields and connected areas. It connects to other guides such as the Glossary of computer science, the Glossary of robotics, the Glossary of machine vision, and the Glossary of logic. These resources help explain the many layers of AI and how they relate to one another.

    A

    A* search is a pathfinding algorithm pronounced "A-star" that guarantees the best solution by never overestimating cost to reach a goal. Abductive logic programming (ALP) lets systems solve problems using incomplete information through abductive reasoning, seeking the most likely explanation for an observation. Ablation studies remove parts of AI systems to see how much each contributes to performance. An abstract data type describes data by its behavior rather than implementation, while abstraction strips away details to focus on what matters. Action languages specify how actions change a system over time and are used in robotics and planning. Adaptive algorithms adjust behavior during use based on rewards or criteria, including the adaptive neuro fuzzy inference system (ANFIS), which blends neural networks with fuzzy logic. A heuristic is admissible if it never overestimates cost, and affective computing studies systems that recognize and simulate human emotion. AI-complete problems are those so difficult they're equivalent to creating truly intelligent machines. AI accelerators are hardware designed to speed up AI tasks like training neural networks, and AI data centers are optimized facilities for running these demanding computations. Algorithms are step-by-step instructions for solving problems, and they're fun.

    B

    Backpropagation trains neural networks by moving error gradients backward through layers, short for "backward propagation of errors," and is key to deep learning using networks with more than one hidden layer. Related techniques include backpropagation through time (BPTT) for recurrent networks like Elman networks, and backpropagation through structure (BPTS) proposed by Christoph Goller and Andreas Küchler in 1996. Backward chaining works from goals backward and is used in theorem provers and AI systems. The bag-of-words model simplifies text as word collections, ignoring grammar and order, used in document classification and computer vision where visual features are treated like words in a vocabulary. Batch normalization, introduced in 2015, stabilizes neural networks by normalizing inputs to each layer. Bayesian programming is a method for probabilistic modeling with incomplete information. The bees algorithm mimics honey bee foraging behavior for optimization problems. Behavior informatics studies behaviors for insights, while behavior trees manage complex task execution in robotics and games. The belief–desire–intention model (BDI) helps agents balance planning and executing tasks. The bias–variance tradeoff describes how models with lower bias tend to have higher variance, and vice versa. Big data refers to datasets too large or complex for traditional tools, often defined by volume, velocity, and variety. Big O notation describes function behavior as it approaches a limit, and binary trees are structures where each node has at most two children.

    C

    A capsule neural network (CapsNet) models hierarchical relationships like biological brains. Case-based reasoning (CBR) solves new problems by referencing past similar cases. Chatbots are computer programs that converse, also called smartbots or conversational interfaces. Cloud robotics connects robots to internet for cloud computing power, making them smarter and cheaper. Cluster analysis groups objects into clusters based on similarity, used in machine learning and image processing. COBWEB, created by Professor Douglas H. Fisher, is an incremental system organizing observations into tree structures for classification. Cognitive architecture describes fixed structures supporting intelligent behavior in natural and artificial systems. Cognitive computing mimics human brain functions to improve decision-making. Commonsense knowledge includes everyday facts like "lemons are sour," first addressed by John McCarthy's 1959 Advice Taker program. Commonsense reasoning simulates how people make assumptions about ordinary situations. Computational creativity combines AI, psychology, philosophy, and arts to explore artificial innovation. Computational linguistics uses computers to study natural language from computational perspectives.

    D

    A computer go program called Darkforest, developed by Facebook, used deep learning and convolutional neural networks. Its updated version, Darkforest2, added Monte Carlo tree search, which mixes tree search methods from chess programs with randomness, becoming known as Darkfmcts3. In 1956, a summer workshop at Dartmouth was seen by many as the start of artificial intelligence as a field. Data augmentation increases data amounts to reduce overfitting in learning algorithms. Data fusion combines multiple sources for more accurate information. Data integration merges data from different places into one unified view, important in both business and science. Data mining finds patterns in large sets using machine learning and statistics. Data science uses methods from math, stats, and computer science to extract insights from all types of data. A dataset is a collection of data, often stored in tables with rows and columns. A data warehouse stores integrated data from various sources for reporting and analysis. Datalog is a logic language used for deductive databases and other applications. Decision boundaries show how neural networks separate data, with more layers allowing more complex shapes. A decision support system helps managers make decisions about unstructured problems. Decision theory studies how agents choose, split into normative and descriptive branches. Decision tree learning uses trees to predict outcomes based on observations. Declarative programming describes what a computation should do without detailing how. A deductive classifier uses frame language to infer from declarations in domains like medicine. Deep Blue was IBM's chess computer that beat a world champion. Deep learning uses neural networks for classification, regression, and representation learning.

    E

    Eager learning builds general rules during training, unlike lazy learning that waits until queries are made. Early stopping prevents overfitting by halting training when performance plateaus. The Ebert test challenges synthesized voices to tell jokes well enough to make people laugh—proposed by Roger Ebert in 2011. An echo state network uses fixed hidden layers with learnable output weights to reproduce patterns over time. Embodied agents interact through physical forms, whether real or virtual. Embodied cognitive science studies intelligence by combining mind and body in holistic models. Error-driven learning aims to reduce feedback errors, a form of reinforcement learning. Ensemble learning combines multiple algorithms for better results. One epoch means training through the full dataset once. The ethics of AI deals with moral issues specific to artificial intelligence. Evolutionary algorithms mimic biological evolution using selection, mutation, and recombination. Evolutionary computation includes these population-based optimization methods. Evolving classification functions handle dynamic data streams. Some fear that advancing artificial general intelligence could lead to existential risks. Expert systems mimic human experts by using rule-based reasoning to solve complex problems.

    F

    A fast-and-frugal tree is a decision-making tool helping models sort data into categories and pick actions based on those groups. A feature is any measurable trait used by machines to understand information—like image edges or user behavior in predictions. Feature extraction takes raw data and builds meaningful characteristics from it, while feature learning lets systems discover these traits automatically instead of relying on hand-crafted inputs. Feature selection whittles down the most useful variables for a model. Federated learning trains models across many devices without centralizing private data. First-order logic uses quantifiers and variables to express ideas more flexibly than basic true/false statements, used in AI reasoning about actions through fluents—conditions that change over time. Formal languages follow strict rules to form valid expressions, and forward chaining moves from known facts toward goals using logical steps. Frames store structured knowledge about stereotyped situations, while frame languages organize this information into hierarchies. The frame problem deals with describing robot environments effectively. Friendly AI aims to build artificial intelligence that benefits humanity, and fuzzy logic allows for degrees of truth between complete falsehood and absolute truth.

    G

    Game theory studies strategic interactions between rational decision-makers, while general game playing aims to design AI that can successfully play multiple games. Generalization lets learners apply past knowledge to new but similar situations, and generalization error measures how well a model predicts unseen data. Generative adversarial networks pit two neural networks against each other in a zero-sum game, and generative AI creates new content by learning patterns from training data. A generative pretrained transformer like GPT first learns to predict the next word in text, then generates human-like responses after being fine-tuned. Genetic algorithms mimic natural selection using operators like mutation and crossover, while glowworm swarm optimization is based on firefly behavior. Gradient boosting improves machine learning by working with pseudo-residuals, and graph theory studies mathematical structures showing relationships between objects. Graph traversal is the process of visiting every vertex in a graph, and graph databases store data as nodes and edges to make relationship-based queries fast and visual.

  5. 05 Artificial intelligence 7m Download (3.2 MB)
    Read this chapter

    Overview

    Artificial intelligence, or AI, is when machines are programmed to do things that we associate with human thinking—like learning, solving problems, and making decisions. Researchers in engineering, math, and computer science work on ways for machines to understand their surroundings and act to reach goals. AI shows up in search engines, chatbots, self-driving cars, game-playing systems like those for chess or Go, and even in creating images or videos. The field started in 1956, and although it had ups and downs over the decades, it really took off after 2012 when new technology like GPUs helped improve neural networks. By 2017, the transformer model pushed progress even further. In the 2020s, generative AI became common, leading to widespread use of tools that can create or change media. As these systems grow more powerful, concerns about safety, ethics, environmental impact, and long-term risks have also increased. Some companies like OpenAI, Google DeepMind, and Meta are working toward artificial general intelligence—systems that could perform almost any cognitive task as well as a human can.

    Goals

    The broad challenge of mimicking intelligence has been divided into smaller tasks that researchers focus on. These tasks represent specific abilities or features that scientists believe an intelligent system should have. The most studied of these traits form the core of current AI research, shaping how experts approach building systems that can think and act like humans. This breakdown helps guide efforts in developing artificial intelligence, making the overall goal more manageable by tackling one capability at a time.

    Reasoning and problem-solving

    In the late 1980s and 1990s, researchers created algorithms that mimicked how people solve puzzles or make logical deductions step by step. These methods were later expanded to handle uncertain or incomplete information using ideas from probability and economics. But as problems grew larger, these approaches suffered from what’s called a "combinatorial explosion"—they slowed down exponentially. Even humans don’t rely on this kind of detailed reasoning all the time; most of the time we use fast, intuitive judgment. Then, in 2024, a new kind of large language model emerged—reasoning models trained to show their intermediate thinking steps. These models improved performance on hard math and coding tasks, though they sometimes produce incorrect outputs or “hallucinations,” unlike older symbolic systems.

    Knowledge representation

    AI systems rely on knowledge to answer questions and draw conclusions about the world. Formal methods of knowledge representation use symbols to stand for words, ideas, and objects, with a knowledge base storing that information in a way a program can access. An ontology outlines the key elements and relationships within a specific field. Researchers have explored these symbolic approaches since the 1970s, but they face challenges—especially with commonsense knowledge, which is vast and often not expressed as clear facts. Another issue is acquiring this knowledge for AI. In contrast, large language models developed since 2012 don't depend on explicit symbols; instead, they learn from enormous amounts of text, like millions of books and billions of webpages. Some AIs also gain understanding through experience, such as AlphaZero mastering games by playing against itself. Though machine learning helps with broad knowledge and common sense, it still struggles with accurate recall and sound reasoning.

    Planning and decision-making

    An agent takes in information and acts on it, whether artificial or not. A rational agent has goals and chooses actions to achieve them. In automated planning, the goal is specific, while in decision-making, agents weigh situations based on preference, assigning utility values. For every action, agents calculate expected utility by considering all possible outcomes and their likelihood, then pick the action with highest expected utility. Classical planning assumes certainty, but real-world problems involve uncertainty—agents must make probabilistic guesses and reassess after acting. Trust comes from being able to explain decisions, especially when they matter. Preferences may be unclear, particularly with humans or other agents, and can be learned or refined through information. The space of possible actions and outcomes is usually too vast to compute fully, so agents must act and evaluate while uncertain. A Markov decision process uses a transition model to predict how actions change states and a reward function to assign utilities and action costs. A policy maps each state to a decision, which can be calculated, heuristically determined, or learned. Game theory helps model rational behavior among interacting agents in AI systems.

    Natural language processing

    Natural language processing, or NLP, is how computers learn to understand and produce human language. Early efforts relied on Noam Chomsky’s generative grammar and semantic networks, but ran into trouble with something called word-sense disambiguation—unless they were limited to small, controlled environments known as “micro-worlds.” British linguist Margaret Masterman argued that meaning, not grammar, was key, and that dictionaries and thesauri should form the basis of computational language structure. Today’s NLP uses deep learning methods like word embeddings and transformers, with attention mechanisms. In 2019, models called GPT began generating coherent text. By 2023, they were scoring at human levels on exams like the bar exam, SAT, and GRE.

    Perception

    Machine perception is how computers make sense of the world using input from sensors like cameras, microphones, and radar. It’s the ability to take in data and figure out what’s happening around them. Computer vision is one part of this, focused on analyzing what’s seen. This field covers things like speech recognition, identifying images, recognizing faces, spotting objects, tracking them as they move, and helping robots understand their surroundings. All of it comes together to let machines interpret the world just like we do, but through technology instead of biology.

    Social intelligence

    Affective computing is the field that deals with systems designed to recognize, interpret, or simulate human emotions. Some virtual assistants are built to sound conversational or even joke around, making them seem more aware of how people feel during a chat, which helps improve how humans interact with computers. But this can mislead users into thinking these systems are smarter than they really are. There have been some modest achievements in this area, like analyzing the sentiment in text, and more recently, combining video and audio to understand emotional expressions better.

  6. 06 Turing test 6m Download (3 MB)
    Read this chapter

    Overview

    The Turing test, created by Alan Turing, was meant to judge whether a machine could act intelligently by mimicking human conversation. In the version still used today, a person evaluates text exchanges between a human and a machine, trying to guess which is which. If the machine’s answers are close enough to human ones, it passes. Turing introduced this idea in his 1950 paper "Computing Machinery and Intelligence" while at the University of Manchester. He called it the imitation game, imagining an interrogator trying to identify a man and a woman in separate rooms. Turing believed that if we can answer whether digital computers could succeed in this game, we’ve answered whether machines can think. Since then, the test has sparked debate, even drawing criticism from philosophers like John Searle.

    Philosophical background

    The question of whether machines can think has deep roots in the debate between dualist and materialist views of the mind. René Descartes, in his 1637 Discourse on the Method, imagined machines that could mimic human speech and reactions, but argued they could never arrange their responses appropriately to any conversation, something even the lowest human can do. This idea prefigures the Turing test, though Descartes did not propose it directly. Later, in 1746, Denis Diderot suggested that if a parrot could answer everything, he’d consider it intelligent—though implicitly limiting the test to natural beings. In 1936, philosopher Alfred Ayer proposed a test for consciousness, saying we can only know if something is truly conscious by whether it passes empirical tests, a notion very similar to Turing’s later approach.

    Cultural background

    In Jonathan Swift's 1726 novel Gulliver's Travels, the king of Brobdingnag initially thinks Gulliver might be a clockwork figure, a marvel of engineering. Even after hearing him speak, the king doubts whether Gulliver was simply taught words to deceive. It's only after asking several questions and receiving thoughtful answers that the king becomes convinced Gulliver isn't a machine. By the 1940s, this kind of test—where humans judge if a computer or alien is smart—was already common in science fiction. Stanley G. Weinbaum's A Martian Odyssey, published in 1934, shows how such tests could be nuanced. Earlier stories also feature beings pretending to be human: the Greek myth of Pygmalion, Carlo Collodi's Pinocchio, and E. T. A. Hoffmann's "The Sandman." In each tale, a human is convinced by something that mimics humanity, at least for a time.

    Alan Turing and the imitation game

    In the 1940s, British researchers including Alan Turing explored machine intelligence, with Turing himself using the term "computer intelligence" as early as 1947. In his 1950 paper “Computing Machinery and Intelligence,” Turing asked whether machines could think, but instead of defining thought, he proposed replacing the question with one about behavior. He introduced the “imitation game,” a version of a party trick where a person and a machine are tested by a judge trying to determine which is which through written questions. Turing imagined a scenario where a machine plays the role of one of the humans, and the judge fails to tell them apart more often than when guessing between a man and a woman. He later refined this idea in a 1952 BBC broadcast, describing a jury questioning a computer meant to convince them it was human. Turing also considered nine objections to machine intelligence, many still debated today.

    The Chinese room

    In 1980, John Searle introduced the "Chinese room" thought experiment in his paper Minds, Brains, and Programs to argue that the Turing test couldn't prove a machine could truly think. He pointed out that programs like ELIZA might pass the test by just manipulating symbols without understanding them. Without real comprehension, Searle believed such systems couldn’t be said to "think" like humans do. His argument sparked intense debate about intelligence, consciousness in machines, and the value of the Turing test through the 1980s and 1990s.

    Loebner Prize

    The Loebner Prize, now reported as defunct, held its first Turing test competition in November 1991, organized by the Cambridge Center for Behavioral Studies in Massachusetts and funded by Hugh Loebner. The contest aimed to push AI research forward, partly because, as Loebner said, no one had acted on the idea of testing machines against humans despite decades of discussion. In 1991, a simple program fooled unsophisticated judges into mistaking it for a human, showing flaws in the test itself. Artificial Linguistic Internet Computer Entity (A.L.I.C.E.) earned bronze honors three times in recent years, and Jabberwacky won in 2005 and 2006. The competition allowed only single-topic conversations at first, but that rule was lifted by 1995. Interactions varied from five minutes to over twenty minutes, and the final event took place in 2019 after funding dried up following Loebner’s death in 2016.

    CAPTCHA

    CAPTCHA systems, designed to distinguish humans from bots online, trace back to early ideas in artificial intelligence and rely on the principles of the Turing test. One widely used version, reCaptcha, was developed by Google. The original versions asked users to identify distorted text or match images, tasks that confuse automated programs. Later, reCaptcha v3 changed the approach by working invisibly in the background, activating when pages load or buttons are clicked, without showing any challenges to users. This method filters out simple bots without interrupting the user experience.

    Attempts

    In 1966, Joseph Weizenbaum built a program named ELIZA that acted like a Rogerian psychotherapist, tricking some users into thinking they were chatting with a real person by repeating keywords from their messages. By 1972, Kenneth Colby had created PARRY, designed to simulate paranoid schizophrenia, which confused judges about half the time when comparing its outputs to actual patient conversations. Then, in 2001, programmers introduced Eugene Goostman, a chatbot that pretended to be a young teenager from Odesa with limited English skills, and in tests, 33% of evaluators believed it was human. Each of these early systems succeeded not through real understanding but by crafting scenarios where their lack of knowledge was excused or ignored.

  7. 07 Guessing 3m Download (1.6 MB)
    Read this chapter

    Overview

    Guessing is how we quickly form a conclusion from what’s right in front of us—something we hold as possible or likely, even though we don’t have enough proof for certainty. A guess isn’t fixed; it's always open to change and depends on what we already know. Often, people use the word without really defining it, assuming it’s clear. But guessing can involve different ways of thinking: deduction, induction, abduction, or even just picking randomly from a set of choices. Sometimes it comes down to intuition—a gut feeling—where someone just knows something is true, even if they can’t explain why.

    Gradations

    Philosopher Mark Tschaepe distinguishes gradations of guessing from wild guesses with no basis to educated ones shaped by prior knowledge. He defines guessing as a deliberate act of creating or selecting potential answers when information is lacking, not merely forming a hunch or stumbling on an answer without reasoning. Tschaepe rejects definitions calling guessing "random" or "instantaneous," noting that what seems unreasoned may actually involve rapid mental processes, as Leibniz observed. A coin flip guess isn't really knowledge but a kind of random selection, unlike complex processes in games like Twenty Questions where reasoning and clues guide guesses. An informed guess uses existing knowledge to narrow possibilities—an educated guess is better than a wild one, as Daniel Wueste noted. Tschaepe also points out that even lucky guesses, which may seem to involve no skill, often rest on suppositions that turn out correct, as William Whewell said about scientific discoveries described as "happy guesses."

    Uses

    Guessing plays a key role in science, especially in forming hypotheses, as Tschaepe notes, calling it an essential part of scientific processes. He links guessing to abductive reasoning, where new ideas are first suggested, describing it as "a combination of musing and logical analysis." Children learn to guess early, often having no other strategy, and develop the ability to recognize when guessing is reasonable even if it's imprecise. In some exams, guessing is penalized, but eliminating wrong answers can make guessing advantageous. Polanyi sees guessing as the outcome of problem-solving, involving clues and direction toward a solution, with a clear process behind it. In literary theory, guessing is necessary because a reader can never fully know an author's intent, so interpreting a text is inherently a guess.

    Software tests

    In software testing, there’s a method called error guessing where testers look for bugs based on what they’ve seen before. They use their experience and gut feeling to decide what might go wrong, drawing from past failures or unexpected issues found during testing. Common problems include things like dividing by zero, null pointers, or bad inputs. There aren’t strict rules for how to do this kind of testing—testers adapt based on the situation, whether that’s using official documents or spotting something strange while working with the software.

    Social impact

    In social settings, guessing can have surprising effects on how people feel about themselves. A study looked at situations where someone guesses another person’s test score or potential salary. It found that sometimes it helps to guess either higher or lower than the actual number. For example, students who knew their own test results were happier when someone else guessed a lower score. That lower guess made them feel like they had done better than expected.

  8. 08 Creativity 6m Download (2.9 MB)
    Read this chapter

    Overview

    Creativity is about coming up with new and useful ideas or works using imagination. These creations can be intangible, like theories or jokes, or physical, like inventions or paintings. It’s also about solving problems in fresh ways. Ancient cultures, such as those in Greece, China, and India, didn’t have a word for creativity—art was seen as discovery, not creation. In the Judeo-Christian-Islamic tradition, creativity was believed to belong only to God, with human efforts seen as reflections of divine work. The modern idea of creativity developed during the Renaissance, shaped by humanist thinking. Today, researchers in psychology, business, and cognitive science study it, along with educators and scholars in the humanities.

    Etymology

    The word “creativity” comes from Latin, where “creare” means “to create,” and its roots trace back even further to “crescere,” meaning “to let things grow.” This idea of growth is central in many indigenous and Eastern views of creativity. The English word "create" first appeared in the 14th century, notably in Chaucer's The Parson's Tale, where it referred to divine creation. It wasn’t until after the Age of Enlightenment that the term came to describe human creative effort.

    Definition

    Creativity, as psychology professor Michael Mumford explained, is about producing something novel and useful, a view shared by Robert Sternberg who said it results in "something original and worthwhile." But over a hundred definitions exist, with Dr. E. Paul Torrance describing it as sensing problems, forming hypotheses, testing them, and communicating findings. Philosophy professor Ignacio L. Götz pointed out that creativity isn't about making something new but about the act of creating itself, noting one can be creative without being original. Creativity differs from innovation, which requires implementation—according to Teresa Amabile and Michael Pratt, or the OECD and Eurostat, who say innovation demands putting ideas into use. There's also emotional creativity, a pattern of mental processes and traits tied to originality and appropriateness in how we feel. From an interdisciplinary angle, creativity strengthens brain connections, improves coherence, and builds social bonds, offering help in coping with despair, hate, and violence.

    Ancient

    Most ancient cultures including Greece, China, and India lacked our modern concept of creativity or a creator. The ancient Greeks used poiein—meaning to make—but only for poetry or poets as makers. Plato questioned whether painters could truly make anything, answering they couldn't because they merely imitated. The idea that humans could create something new wasn't part of Western thought until later. Scholars say the modern concept came from Christianity, rooted in Genesis's creation story. But even then, creativity belonged solely to God in Judeo-Christian-Islamic tradition. Humans were seen as vessels for divine inspiration, not creators themselves. Greeks and Romans believed in external forces like the Muses or divine daemon/genius that inspired art, but this wasn't creativity as we know it now. These ideas dominated Western thinking until the Renaissance.

    Renaissance

    Creativity began to be seen differently during the Renaissance, not as a gift from the divine, but as something that came from human skill and talent. This shift was driven by humanism, a movement that placed people at the center of attention, celebrating individual intellect and achievement. It was during this time that the idea of the "Renaissance man" emerged — someone who pursued knowledge and creativity across many fields. Leonardo da Vinci stands as one of the most famous examples of such a person, embodying the spirit of this era's ideals.

    From the 17th to the 19th centuries

    The idea that creativity comes from individual ability rather than divine inspiration developed slowly, becoming clear during the Age of Enlightenment. By the 18th century, people began linking creativity with imagination, especially in art. Thomas Hobbes saw imagination as central to how humans think, and William Duff was among the first to call it a mark of genius, distinguishing it from mere talent. It wasn’t until the 19th century that creativity started being studied seriously. Psychologists Mark Runco and Robert Albert point to the late 1800s, when Darwinism sparked interest in individual differences. Francis Galton, influenced by eugenics, studied intelligence and creativity as parts of genius.

    Modern

    In the late 1800s and early 1900s, thinkers like Hermann von Helmholtz and Henri Poincaré reflected on creativity, inspiring later theorists such as Graham Wallas and Max Wertheimer. Wallas outlined a five-stage model of creativity in his 1926 book Art of Thought, which includes preparation, incubation, intimation, illumination, and verification. He saw creativity as an evolutionary tool for adapting to change, a view updated by Simonton in Origins of Genius. In 1927, Alfred North Whitehead coined the term “creativity” in his Process and Reality, while early psychometric studies of imagination were done by H.L. Hargreaves at the London School of Psychology. The formal study of creativity as a distinct cognitive ability began with J.P. Guilford’s 1950 address to the American Psychological Association, which showed that creativity isn’t tied to IQ above a certain level.

    Across cultures

    Creativity means different things across cultures. In Hong Kong, researchers found that Westerners see it through individual traits like aesthetic taste, while Chinese people focus more on how creativity benefits society. Mpofu and colleagues looked at 28 African languages and discovered that 27 lacked a direct translation for “creativity,” with Arabic being the only exception. This connects to the linguistic relativity hypothesis, which suggests language can shape thought, though more study is needed to confirm that. There's also been little research on creativity in Africa and Latin America, even though it’s been studied more in the northern hemisphere. Even within the north, views vary—Scandinavians tend to see creativity as a personal tool for coping, while Germans view it more as a problem-solving method.

Read

Free to download, keep and share. For general information only — not professional medical, legal or financial advice. Please consult a qualified professional.

← All audiobooks