The Efficiency Frontier in Financial Intelligence Software developers often default to a bigger-is-better mentality when Large Language Models fail. If a model cannot solve a complex financial query, the standard industry response is to swap it for a massive, high-parameter alternative. However, Kobie Crawford from Snorkel argues that this "sledgehammer to crack a walnut" approach is both inefficient and unnecessary for many enterprise tasks. By focusing on behavior rather than raw reasoning depth, developers can achieve elite performance from smaller, faster models. Solving the Terence Tao Effect The research, a collaboration between Snorkel and the RLLM team at UC Berkeley, highlights what researchers call the Terence Tao effect. Much like a world-class mathematician who can solve any abstract proof but might struggle with a specific accounting database, massive models like Qwen 3 235B possess immense reasoning power but lack discipline in tool execution. When tasked with analyzing YouTube ad revenue, the 235B model bypassed environment inspection entirely, queried non-existent tables, and eventually hallucinated an answer. It wasn't a lack of intelligence; it was a lack of behavioral constraint. Engineering Discipline Through GRPO To bridge this gap, the team utilized Group Relative Policy Optimization (GRPO) to fine-tune a tiny 4B parameter model. The objective was simple: teach the model how to interact with its environment before attempting to answer a question. Using a FinQA environment, the model learned a specific sequence: call `get_table_names`, inspect the schema, run the query, and self-correct if a column error occurs. This behavioral shift allowed the smaller model to succeed where the giant failed, transforming it into a reliable financial agent. The Paradox of Simple Training Data One of the most striking findings from the Snorkel study was the efficacy of curriculum learning. Researchers compared training regimes using single-table questions, multi-table questions, and a mixture of both. Surprisingly, training exclusively on single-table tasks yielded the highest performance gains. This "single-step" focus fixed the core failure mode of tool use so effectively that the improvements generalized to the more difficult FinQA Reasoning benchmark, where the model's accuracy doubled from 13.9% to 26.6%. Rubrics Over Binary Feedback For developers looking to replicate these results, Kobie Crawford emphasizes the importance of evaluation rubrics. Instead of a simple pass/fail metric, a detailed rubric breaks down the model's response into specific components: Did it check the table names? Did it verify the schema? By identifying the exact point of failure, developers can generate high-quality, expert-verified data sets that target specific behavioral weaknesses. This methodical approach ensures that Reinforcement Learning (RL) cycles, which cost less than $500 per run in this study, remain both tractable and highly effective for production-grade software development.
Large Language Models
Products
Jul 2023 • 1 videos
Lighter month. ArjanCodes covered Large Language Models across 1 videos.
Jun 2025 • 1 videos
Lighter month. The Riding Unicorns Podcast covered Large Language Models across 1 videos.
Aug 2025 • 1 videos
Lighter month. Chris Williamson covered Large Language Models across 1 videos.
Sep 2025 • 2 videos
High activity month for Large Language Models. ArjanCodes and The Riding Unicorns Podcast among the most active voices, with 2 videos across 2 sources.
Nov 2025 • 1 videos
Lighter month. AI Engineer covered Large Language Models across 1 videos.
Dec 2025 • 1 videos
Lighter month. ArjanCodes covered Large Language Models across 1 videos.
Jan 2026 • 2 videos
High activity month for Large Language Models. Laravel Daily and AI Engineer among the most active voices, with 2 videos across 2 sources.
Mar 2026 • 2 videos
High activity month for Large Language Models. Chris Williamson and The Prof G Pod – Scott Galloway among the most active voices, with 2 videos across 2 sources.
Apr 2026 • 2 videos
High activity month for Large Language Models. The Iced Coffee Hour and The Prof G Pod – Scott Galloway among the most active voices, with 2 videos across 2 sources.
Jun 2026 • 2 videos
High activity month for Large Language Models. AI Engineer among the most active voices, with 2 videos across 1 sources.
ArjanCodes (2 mentions) discusses Large Language Models in the context of resilient Python code and AI agent design, while Laravel Daily mentions using Large Language Models to generate Laravel migrations.
- Jun 10, 2026
- Jun 1, 2026
- Apr 19, 2026
- Apr 8, 2026
- Mar 30, 2026
The Hidden Tax of the Hyperactive Hive Mind Ten years after Cal Newport released his seminal work on concentration, the state of the modern workplace has arguably regressed. We are currently caught in the gravitational pull of what Newport calls the hyperactive hive mind—a style of collaboration defined by ad hoc, unscheduled communication that demands constant attention. This environment isn't just a nuisance; it is a fundamental mismatch for the human brain's evolutionary hardware. Our minds require significant time to transition between abstract symbolic tasks, yet data from Microsoft 365 reveals that the average knowledge worker now switches context every two minutes. This constant ping-pong match of Slack messages and Microsoft Teams notifications creates a state of diffuse cognitive friction. When we are interrupted mid-thought, it takes roughly ten to twenty minutes for our brains to fully load the relevant information for a new task. If we are interrupted every two minutes, we never truly "lock in." The result is a workforce that is perpetually fatigued, spending their weekdays talking about work while pushing the actual high-value output—the "deep work"—to Saturday and Sunday mornings when the digital noise finally subsides. This is a massive economic failure, representing a remarkably low return on the high-priced human brains companies employ. Why AI Work Slop Is Making Us Dumber The arrival of large language models like ChatGPT was initially hailed as a productivity savior, but it has introduced a new toxin: work slop. This term describes AI-generated reports, emails, and presentations that are low in quality but high in volume. Because our brains are already exhausted by the hyperactive hive mind, we are increasingly using AI to avoid the painful spikes of peak concentration. We ask the machine to fill the blank page, resulting in wordy, vacuous documents that make everyone else's job harder by forcing them to sift through noise to find the signal. Cal Newport argues that this creates a dangerous feedback loop. We are already primed to dislike heavy cognitive load, and our comfort with concentration has been further degraded by algorithmic distraction machines like TikTok. When AI offers a way to smooth over the peaks of cognitive strain, we take it. However, the market ultimately pays for economic value, not busyness. AI-generated work slop doesn't generate value; it creates administrative overhead. The real competitive advantage in the coming years will not belong to those who can prompt an LLM to write an email, but to those who maintain the rare ability to tolerate cognitive strain and produce original, high-quality work. The Kaplan Curve and the LLM Asymptote There is a prevailing belief that AI will continue to improve at an exponential rate until it achieves Artificial General Intelligence (AGI). This belief stems from the Kaplan Curve, a 2020 observation that increasing the size and training time of LLMs lead to predictable performance gains. This held true from GPT-2 to GPT-4, the latter of which began showing surprising logical and mathematical abilities. However, newer projects like OpenAI's Orion and Meta's Behemoth are reportedly hitting a brick wall. Simply making models bigger is no longer yielding the same dramatic leaps in capability. We are likely reaching an asymptote for pure transformer-based architectures. The future of AI will likely shift from giant, general-purpose oracles to distributed, bespoke systems. These hybrid models will combine LLMs with explicit logic engines and world models designed for specific tasks—such as an AI that plays chess better than a human versus one that manages customer service. For the individual, this means that while certain narrow fields will be automated, the dream of a singular "god in a box" that replaces all human cognition is receding. The need for human experts who can manage these complex tools and provide the "last mile" of high-resolution thinking is actually increasing. Rebuilding the Individual Capacity for Focus To thrive in this landscape, we must treat focus as a tier-one skill rather than a personality trait. Cal Newport suggests that reading physical books is the cognitive equivalent of "getting your steps in." The process of reading long-form text rewired the human brain during the Neolithical revolution, yoking together disparate parts of the brain to process sophisticated thoughts. When we read exclusively on screens, we tend to skim and jump, which keeps our thinking shallow. Physical books—or Kindle devices that mimic the physical page—force us to spend time under tension with complex ideas. Furthermore, we must change our relationship with cognitive strain. Athletes understand that the burn of a muscle signifies growth; knowledge workers must learn to view the "itch" of boredom or the difficulty of a complex problem as the feeling of their brain becoming more capable. While the rest of the world uses AI to run away from strain, those who run toward it will become the superstars of the knowledge economy. You cannot hide behind busyness forever because busyness cannot be monetized. If you produce rare and valuable things, you gain the leverage to write your own ticket—exempting yourself from the meetings and digital clutter that define the average corporate existence. Rescuing the Organization from the Local Minimum At the organizational level, the hyperactive hive mind persists because it is the "low energy state" of work. It requires the least amount of planning and structure, even though it is wildly inefficient. To escape this trap, leaders must implement explicit workload tracking. No one should simply have tasks "thrown" at them. Instead, projects should live in a team-wide queue, and individuals should only pull three or four things onto their personal plate at a time. Once a task is assigned, it generates an "administrative tax" of emails and meetings; by limiting work-in-progress, you drastically reduce this overhead. Finally, organizations must kill the expectation of constant accessibility. Newport proposes a rule: if a message requires more than one response, it must happen in real-time. This can be managed through daily office hours or morning stand-ups where teams coordinate their needs for the day in ten minutes, rather than letting a ping-pong match of Slack messages unfold over five hours. When you make people accountable for their output rather than their responsiveness, you transform the culture. In an era where AI can automate the mundane, the ultimate organizational asset is a team that has the time and the silence to actually think.
Mar 5, 2026Overview Transitioning from a mock backend to a production-ready API shouldn't require a total architectural overhaul. This guide demonstrates how React and Next.js developers can swap a local JSON Server for a robust Laravel backend with minimal friction. By adhering to standard RESTful conventions, you can achieve persistence and scalability without rewriting your frontend data-fetching logic. Prerequisites To follow this implementation, you should have a baseline understanding of JavaScript (specifically TypeScript interfaces) and PHP fundamentals. Familiarity with the fetch API or Axios is necessary for the frontend, while a basic grasp of relational databases like MySQL will help on the server side. Key Libraries & Tools * **Laravel**: A PHP framework providing the API structure. * **Eloquent ORM**: Laravel's built-in database mapper. * **MySQL**: The relational database for persistent storage. * **Next.js**: The React-based framework for the user interface. Code Walkthrough Frontend Configuration Simply update your base URL. If your React app previously pointed to a local JSON file, redirect it to the Laravel endpoint: ```javascript // From local mock server const BASE_URL = "http://localhost:3000/posts"; // To Laravel API const BASE_URL = "http://api.test/api/posts"; ``` Backend Implementation On the Laravel side, define a resource route in `routes/api.php`. This single line handles index, store, update, and delete requests. ```php Route::apiResource('posts', PostController::class); ``` Next, generate the necessary boilerplate using **Artisan commands**. These automate the creation of the controller, the Eloquent model, and the database migration: ```bash php artisan make:model Post -m -c --api ``` In the `PostController`, the `index` method returns all records as JSON, matching the structure your React components already expect: ```php public function index() { return Post::all(); } ``` Syntax Notes Laravel uses **Route Resources** to map HTTP verbs to controller actions automatically (e.g., `GET /posts` maps to `index()`). Additionally, Eloquent models automatically include `created_at` and `updated_at` timestamps in the JSON response, which provides better data auditing than manual JSON mocks. Practical Examples This setup is ideal for scaling a prototype. You might start with a JSON Server for a "proof of concept" dashboard, then switch to Laravel when you need to implement secure authentication, complex relations, or real-time MySQL data processing. Tips & Gotchas Ensure your Laravel migration file defines every field used in your React frontend, or the API will throw a 500 error during `POST` requests. Use Large Language Models like ChatGPT to generate these migrations quickly, as Laravel’s stable syntax makes it highly predictable for AI assistance.
Jan 22, 2026Beyond the REST wrapper Software development is currently grappling with a fundamental misunderstanding of how artificial intelligence interacts with code. Jeremiah Lowin, founder and CEO of Prefect%20Technologies, argues that the industry is flooding the ecosystem with \"bad\" Model%20Context%20Protocol (MCP) servers. The core of the problem lies in the lazy habit of treating AI%20agents like human developers. Most creators simply point an MCP generator at an existing REST%20API and ship the resulting wrapper, assuming that if a human can use an endpoint, an AI can too. This assumption is not just wrong; it is architecturally damaging. Humans do not actually use APIs in their raw form; we build websites, mobile apps, and SDKs to shield ourselves from them. Agents, however, are forced to interface directly with these programmatic structures. When we give an agent a raw REST API, we are handing it a toolset designed for human strengths—one-time discovery and fast, cheap iteration—while ignoring the reality of agentic limitations. To build effective MCP servers, we must stop thinking in terms of infrastructure and start thinking in terms of product design. This means moving beyond the transport layer and focusing on the user interface of the future: the agentic interface. Three pillars of agentic interaction Designing for agents requires a fundamental shift in how we view discovery, iteration, and context. For a human developer, discovery is a one-time cost. You read the documentation once, understand the schema, and never look at the docs again for the life of that application. For an agent, discovery happens every single time it handshakes with a server. It must enumerate every tool and description, a process that is incredibly expensive in terms of token consumption and latency. Iteration follows a similar pattern of divergence. Human developers can run a script dozens of times a minute to debug a connection; it is fast and essentially free. For an agent, every additional call involves sending the entire history of previous calls back over the wire. Iteration is the enemy of performance. Finally, we must respect the limitations of context. While a human draws on years of memory and experience, even the most advanced Large%20Language%20Models (LLMs) operate with a limited \"brain\" of roughly 200,000 tokens. If your MCP server lobotomizes the agent on the initial handshake by consuming half its context window with documentation, the agent will fail before it even starts. The goal of a great MCP server is to find the needle in the haystack without forcing the agent to inspect every single piece of hay. Outcomes over atomic operations Engineers are traditionally trained to build atomic, granular operations. In a REST API, having separate endpoints for `get_user`, `get_orders`, and `filter_results` is considered a best practice. In an MCP server, this is a failure. Forcing an agent to act as a glue layer or an orchestrator is slow, expensive, and stochastic. Every round trip increases the chance of a hallucination or a timeout. Instead, we must design for outcomes. A single tool should map to a single agent story. Rather than providing three atomic tools that the agent must stitch together, provide one tool called `track_latest_order_by_email`. This buries the complexity of the API calls within the server-side logic, where it is fast and deterministic, and presents the agent with a clear path to success. This \"top-down\" workflow design ensures that the agent is not wasting its limited reasoning capabilities on basic plumbing that a human developer should have already handled. Argument flattening and schema discipline A recurring trap in MCP development is the use of complex, nested arguments. Developers often pass dictionaries or configuration objects, assuming the agent will intuit the structure. This often leads to \"doubly documented\" tools where the system prompt and the tool schema disagree, leaving the agent to guess which is correct. Lowin points out that many popular clients, including Claude%20Desktop, have historically struggled with structured object arguments, sometimes converting them to strings and breaking the handshake. The solution is radical simplicity: flatten your arguments. Use top-level primitives like strings, integers, and booleans. When choices are constrained, utilize literals or enums to provide a strict boundary for the agent. This reduces the cognitive load on the model and ensures that the arguments sent from the client match the expectations of the server. By naming arguments for the agent rather than the developer, you create an interface that is self-describing and resistant to the common pitfalls of LLM-based tool calling. The economy of the token budget The most overlooked constraint in MCP design is the token budget. Every character in a docstring, every field in a schema, and every example in a prompt costs tokens. Lowin highlights a cautionary tale of a company attempting to expose 800 endpoints via MCP. At that scale, the mere act of telling the agent what tools are available consumes the entire context window, leaving zero room for actual problem-solving. This is the \"handshake overflow,\" and it is the primary reason many enterprise-grade MCP servers feel sluggish or \"dumb.\" We must curate ruthlessly. A server with more than 50 tools is a red flag. If you find yourself exceeding this limit, it is time to namespace your tools or split the server into smaller, specialized units. One of the most effective ways to save tokens is to use \"progressive disclosure.\" Instead of documenting every edge case in the main tool description, use helpful error messages. In the agentic world, errors are prompts. If an agent calls a tool incorrectly, the server should return an error message that provides the necessary context to fix the call. This moves the token cost from the handshake (which happens every time) to the error recovery (which happens only when needed). Curation as the ultimate design goal The transition from infrastructure-thinking to product-thinking is the final hurdle for the MCP ecosystem. Tools like FastMCP have made it trivial to bootstrap a server by mirroring a REST API, but this should only ever be a starting point. The real work begins with curation. We must move toward \"context products\". These are highly optimized, specific interfaces that give an agent exactly what it needs to achieve a specific result, and nothing more. Treating an MCP server as a user interface changes the development lifecycle. It requires evaluation, testing against specific models, and an understanding of how different clients like Cursor or Claude%20Code handle the protocol. As the ecosystem matures, the focus will shift away from how we connect models to data and toward how we design the surface area of those connections. The developers who succeed will be those who view their servers not as a collection of functions, but as a carefully crafted experience for the most important new user in tech: the AI agent.
Jan 12, 2026Overview of the Retry Pattern In modern software development, your code rarely lives in a vacuum. It communicates with databases, external APIs, and Large Language Models (LLMs). These connections are prone to **transient failures**—brief, temporary issues like network hiccups or rate limits that cause a script to crash even when the logic is perfect. The **Retry Pattern** solves this by wrapping potentially flaky operations in a loop that automatically attempts the action again before giving up. This simple architectural shift transforms fragile scripts into robust, production-grade applications. Prerequisites To follow this guide, you should be comfortable with Python fundamentals, specifically: * **Higher-order functions**: Understanding how to pass functions as arguments. * **Decorators**: Familiarity with the `@` syntax and function wrapping. * **Type Hinting**: Knowledge of `Callable`, `Generics`, and the `typing` module. * **Exception Handling**: Using `try/except` blocks to manage errors. Key Libraries & Tools * Python (v3.10+ recommended for advanced typing). * **functools**: A standard library used for `wraps` to maintain function metadata. * Tenacity: A powerful, specialized library for retrying tasks in production. * SerpApi: A tool for reliable search engine data extraction that handles retries internally. Code Walkthrough: From Simple Loops to Decorators 1. The Basic Retry Loop A manual retry function uses a `range` loop to attempt an operation. If the operation succeeds, it returns immediately; if it fails, it sleeps before the next attempt. ```python from typing import Callable, TypeVar import time T = TypeVar("T") def retry(operation: Callable[[], T], retries: int = 3, delay: float = 1.0) -> T: for attempt in range(1, retries + 1): try: return operation() except Exception as e: if attempt == retries: raise e time.sleep(delay) ``` 2. Exponential Backoff Retrying too quickly can overwhelm a struggling server. **Exponential backoff** increases the wait time after each failure, giving the remote service breathing room to recover. ```python Inside the retry logic sleep_time = delay * (backoff_factor ** (attempt - 1)) time.sleep(sleep_time) ``` 3. The Decorator Implementation To make retry logic reusable across your entire codebase without manual calls, we can implement a decorator. This uses `functools.wraps` to ensure that the decorated function retains its original name and docstring. ```python from functools import wraps def retry_decorator(retries=3, delay=1.0): def decorator(func): @wraps(func) def wrapper(*args, **kwargs): for attempt in range(retries): try: return func(*args, **kwargs) except Exception as e: if attempt == retries - 1: raise e time.sleep(delay) return wrapper return decorator ``` Syntax Notes: Callable and Generics When building these utilities, use **TypeVars** (like `T`) to ensure the `retry` function returns the same type as the original operation. This maintains IDE autocomplete and type safety. Additionally, the `Callable[[], T]` syntax specifies that the function takes no arguments and returns type `T`. For functions with arguments, use `Callable[..., T]` or specific parameter lists. Practical Examples * **LLM JSON Parsing**: LLMs like those from OpenAI occasionally return malformed JSON. A retry pattern allows the code to re-prompt or re-parse the response without crashing the pipeline. * **Web Scraping**: When using tools like SerpApi, retries handle network timeouts or rotating proxy shifts automatically, ensuring data consistency. * **Database Connections**: Brief lockouts or connection resets can be mitigated with a 2-second retry window. Tips & Gotchas * **Avoid Permanent Errors**: Never retry a 404 (Not Found) or 401 (Unauthorized) error. Retrying will not fix a wrong URL or an invalid API key; it only wastes resources. * **Side Effects**: Ensure the operation is **idempotent**. If a function writes to a database before failing, retrying might create duplicate entries. * **Production Use**: For mission-critical code, use Tenacity. It offers advanced features like "jitter" (randomized delays) to prevent "retry storms" where multiple clients hit a server simultaneously.
Dec 19, 2025The Myth of the Unavoidable Bug Most users experience software as something that "just works" until it suddenly doesn't. For the person using a banking app or a camera, a bug is a fleeting frustration. For the developer, however, bugs are a source of constant atmospheric pressure—a reality of on-call rotations, pager alerts, and the relentless creep of technical debt. We have conditioned ourselves to believe that perfection is impossible, citing millions of lines of code, ambiguous specifications, and the sheer unpredictability of the physical world. Johann Schleier-Smith from Temporal Technologies challenges this defeatist status quo. He argues that the industry already knows how to build Zero-Bug Software. The methodologies have existed for decades, tucked away in the high-stakes corridors of aerospace and medical engineering. The primary barrier has never been a lack of knowledge; it has been the crushing weight of economics. High-assurance software traditionally costs upwards of $2,500 per line of code, a price point that renders it inaccessible for 99% of commercial applications. We are now entering an era where AI agents could bridge this 100x cost gap, making aerospace-grade reliability the default for every digital interaction. Lessons from the Flight Deck and Deep Space The Airbus A320 stands as a monument to what is possible when the industry rejects defect tolerance. Its control software, developed in the 1980s, has never been implicated in a serious flight incident. This wasn't achieved through luck, but through a rigorous adherence to N-version programming: separate teams using different processors (Intel x86 versus Motorola) and distinct operating systems to ensure that a single logic error couldn't bring down the aircraft. Similarly, NASA demonstrated near-perfection with the Space Shuttle program. Over its final versions, the software averaged only one error per 420,000 lines of code. This level of precision is roughly 1,000 times more reliable than typical commercial software. These systems prioritize static memory allocation, explicit error handling, and the total decoupling of verification teams from development teams. While critics argue that such processes stifle innovation, the data suggests that quality through process is the only proven path to absolute reliability. The Three Pillars of Manageable Complexity To understand how we move toward zero bugs, we must revisit the foundation of computer science. The first pillar is the high-level language. By moving away from assembly in the 1950s and 60s, we gained a 10x productivity boost by abstracting machine implementation details like registers and memory layout. This allows us to focus on "essential complexity"—the logic of the problem itself—rather than the quirks of the hardware. Edgar Dijkstra introduced the second pillar: structured programming. By eliminating the "go-to" statement and replacing it with sequences, selections, and iterations, developers gained the ability to use compositional reasoning. This means you can understand a block of code by looking at its immediate context rather than tracing a tangled web of jumps. Finally, David Parnas gave us modularity. Modularity allows for local reasoning, ensuring that as systems grow, the complexity scales linearly rather than exponentially. These three pillars are not just historical footnotes; they are the exact features that make code interpretable for Large Language Models (LLMs) today. Formal Methods and the Power of Proof While testing only proves the presence of bugs, formal methods can prove their absence. Languages like Daphne allow developers to write proofs directly alongside their code. When you run a verifier, it uses automated reasoning to ensure that every assertion holds true across all possible execution paths. We are seeing a renaissance in these techniques. The seL4 microkernel is a fully verified operating system used in security-critical applications. The CompCert compiler is a verified C compiler that guarantees the generated machine code exactly matches the source program’s intent. Even the Internet itself is increasingly protected by Project Everest, which provides verified cryptographic libraries. The speed and success rates of these verification tools have improved by orders of magnitude over the last 20 years, turning what was once a theoretical academic exercise into a commercially viable toolset. Engineering the Agentic Future The rise of Agentic Coding introduces a paradox. While LLMs are non-deterministic and prone to hallucinations, they possess a unique resilience: the ability to handle ambiguity and unanticipated inputs that would crash traditional rigid software. The key to "Software 3.0"—as Andrej Karpathy calls it—is applying old high-assurance processes to new AI workflows. Instead of asking an LLM to just "write code," we should be prompting it to conduct explicit risk analysis and write "safety cases" for its logic. We can emulate the Airbus model by using one foundation model (like GPT-4) to write the tests and another (Claude) to write the code. When agents are tasked with verifying their own work through formal methods, the cost of high-assurance code plummets. Schleier-Smith notes that while human-written high-assurance code costs $2,500 per line, agent-generated code can be produced for pennies. This 10,000x reduction in cost is the catalyst for the zero-bug vision. Once agents routinely produce software with fewer defects than humans, adoption will reach a point of absolute takeoff, fundamentally altering our expectations of what software can—and should—be.
Nov 24, 2025Overview Modern AI development often begins with simple prompts but quickly devolves into unmanageable "spaghetti code" as developers add tools, data pipelines, and multiple model calls. This tutorial demonstrates how to apply classic software design patterns—Chain of Responsibility, Observer, and Strategy—to Python AI agents. By decoupling logic from prompts, you create modular systems that are easier to test, debug, and scale. These patterns transform one-off hacks into professional, maintainable software architectures. Prerequisites To follow this guide, you should have a solid grasp of Python basics, specifically functions and lists. Familiarity with Pydantic for data validation is helpful, along with a baseline understanding of how Large Language Models (LLMs) interact with system prompts and context. Key Libraries & Tools * **Pydantic AI**: A framework for building production-grade agents with built-in validation. * **OpenAI API**: The underlying model provider for generating agent responses. * **Python `typing` module**: Used for defining protocols and callable types to ensure structural typing. Code Walkthrough The Chain of Responsibility This pattern allows a series of specialized agents to process a request sequentially. Each step handles a specific concern—like finding a hotel or booking a flight—and passes the updated context to the next handler. ```python def plan_trip(user_input, deps): context = TripContext() # The Chain: A list of callables executed in sequence chain = [handle_destination, handle_flight, handle_hotel, handle_activities] for handler in chain: handler(user_input, deps, context) return context ``` The Observer Pattern Monitoring agent behavior is critical for debugging. The Observer pattern allows you to attach logging or monitoring tools without polluting your core business logic. ```python class AgentObserver(Protocol): def notify(self, agent_name: str, prompt: str, duration: float): ... def run_with_observers(agent, prompt, observers): start = time.time() output = agent.run(prompt) duration = time.time() - start for obs in observers: obs.notify(agent.name, prompt, duration) return output ``` The Strategy Pattern Use the Strategy pattern to swap agent behaviors or "personalities" dynamically. Instead of hardcoding prompts, you pass a strategy function that returns a preconfigured agent. ```python def run_travel_strategy(strategy_func, prompt): # The strategy_func acts as a pluggable behavior factory agent = strategy_func() return agent.run(prompt) Usage run_travel_strategy(get_budget_agent, "I need a trip to Paris") ``` Syntax Notes We utilize **Protocols** from Python's `typing` module to implement structural subtyping. This allows us to define what an "Observer" looks like without forcing rigid inheritance hierarchies. Additionally, using **Callables** as strategies keeps the implementation functional and lightweight compared to traditional class-heavy patterns. Practical Examples These patterns excel in multi-step workflows such as automated customer support triaging (Chain), real-time performance dashboards (Observer), or persona-driven marketing copy generation (Strategy). Tips & Gotchas Always ensure your context object is passed by reference through the chain to maintain state. A common mistake is failing to handle errors midway through a chain; if the destination agent fails, the flight agent should likely never run. Implement early exits or error-handling strategies within your loop to prevent cascading failures.
Sep 12, 2025The Death of Artisanal Software and the Rise of the AI Native Founder We are witnessing a fundamental shift in how companies are built, transitioning from a world where humans wrote 80% of code to one where 80% is generated by models. This isn't just a technical evolution; it's an existential change for the startup ecosystem. As a former operator at Microsoft and Stripe, I’ve seen the transition from hand-crafted "artisanal" software to what is now becoming "mass-produced" software. For the first time since the 1960s, the capabilities we once only dreamed of in computer science are becoming reality through Large Language Models. The barrier to entry for prototyping has vanished. We are now in the era of "vibe coding," where a founder with a clear vision can iterate faster than a traditional engineering team ever could. This creates a new expectation in the venture capital world. If you show up to a pitch for a pre-seed or seed round without a working prototype, you are sending a signal that you haven't embraced the current paradigm. AI native founders are prioritizing building over deck-perfecting, and those who spend their nights vibe coding are the ones winning the market. The New Economics of Capital Efficiency and Distribution In the previous generation of startups, a seed round was essentially a hiring mandate. You raised a few million dollars to hire five engineers and sat in a basement for nine months to ship a product. Today, the AI native playbook is radically different. We are seeing founders hire a single engineer and then spend their remaining budget on "fleets of agents," tokens, and sophisticated workflows. The cost of building has collapsed, leading to a massive reallocation of capital toward distribution, brand, and marketing. This capital efficiency is creating a competitive environment where speed is the primary weapon. One of the most striking pitches I've seen recently featured a founding team comprised of an engineering manager and five "Devins" from Cognition AI. For roughly $2,500 a month, they were doing the work that would have previously cost hundreds of thousands in payroll. This shift forces us to rethink what a "company" actually looks like. If the cost of the "act of building" goes to near zero, then value must be found elsewhere. Defensibility in a World of Carbon-Copy Software If an agent can look at a competitor’s website and replicate a feature in an afternoon, where does defensibility come from? The answer lies in the "good old moats" of the 2010s: distribution, data, taste, and brand. To survive, founders must become subject matter experts who own the holistic workflow of a problem. A customer buys Linear not because they can't find another issue tracker, but because the team at Linear has the best "taste" and expertise in how project management should actually work. Owning the workflow is also the only way to build a data moat. By facilitating the full journey of solving a problem, you collect the specific reinforcement learning data needed to train agents that are better than generic models. A generic AI won't know the nuances of a specific accounting operation or how a venture capitalist reviews a deal. If you don't own the workflow, you can't collect the data, and if you can't collect the data, you can't build a specialized agentic system. This is where the next generation of giants will be built. Agent Experience is the New Developer Experience We are moving beyond Customer Experience (CX) and Developer Experience (DX) into the era of Agent Experience (AX). As startups increasingly use tools like Lovable, Cursor, and Replit to build their products, the underlying infrastructure must adapt. These "vibe coding" tools are not just toys; they are the new primary users of APIs. Take Resend as an example. When a user asks Lovable to build an email flow, the agent recommends Resend. This creates a massive growth loop where the GDP of a business is directly correlated to the GDP of vibe coding. Infrastructure providers now need to treat agents as a first-class client type. This means optimizing APIs for agent consumption, much like we once optimized web experiences for mobile phones. My former team at Stripe is already doing this with specialized servers that agents can talk to directly. If you aren't optimizing for agents, you are invisible to the most productive builders in the market. Bridging the Atlantic Gap in Tech Ambition Having spent decades in both Copenhagen and New York, the cultural divide between European and American tech ecosystems remains stark. In Denmark, there is often a "tall poppy" syndrome where success is defined by a stable middle-management role. While this has improved, the US still holds a significant lead in celebrating risk and taking "big swings." Europe has traditionally used American primitives to build vertical SaaS, but the next decade offers an opportunity for Europe to build its own sovereign infrastructure and cloud primitives in a new geopolitical reality. However, for a European founder to truly scale, they must adopt a global mindset early. Expanding from Denmark to Germany isn't a big swing; the real market is the US. New York City has emerged as the ideal landing spot for these founders. It is the second-largest tech ecosystem in the world and offers a time zone that allows for seamless collaboration with engineering teams back in Lisbon, Stockholm, or Copenhagen. If you want to build a foundational company, you need to be where your customers are, and for enterprise tech and AI, that is increasingly New York. Inside the AlleyCorp Incubation Machine At AlleyCorp, we don't just wait for the right founder to walk through the door; we build the companies we want to see. Our incubation process is born from operational conviction. If we see a tangible problem in healthcare, robotics, or AI that nobody is solving correctly, we put a team together and lead as the interim CEO. This allows us to lean into our experience as former operators to de-risk the earliest stages of company building. A prime example is Radical AI. We saw a massive opportunity at the intersection of material science and AI, incubated the team, and a year later they raised $60 million to build foundational models for new materials. This model works because we have an in-house engineering team that acts as an execution capacity for our portfolio. We aren't just writing checks; we are building the machine that builds the companies. In an agentic world, this ability to rapidly prototype and validate ideas is the ultimate competitive advantage.
Sep 10, 2025The Mirror of Machine Intelligence When we look at Artificial Intelligence, we aren't just seeing a tool; we are seeing a reflection of our own cognitive architecture. For centuries, humans have held reasoning as our primary claim to uniqueness. Aristotle believed it was the one thing that separated us from the animals. Yet, our progress in building Large Language Models has revealed a startling inversion of this assumption. This phenomenon, known as Moravec's Paradox, highlights that high-level reasoning and arithmetic—tasks we find difficult—are computationally easy for machines. Meanwhile, the simple act of carrying a cup of water or cracking an egg remains an insurmountable challenge for modern robotics. This discrepancy exists because evolution has spent four billion years optimizing our motor skills and sensory perception. Reasoning and abstract logic are, in evolutionary terms, brand-new software patches developed over only the last million years. By attempting to replicate human ability in silicon, we have discovered that our "primal" abilities are actually our most sophisticated. We are now in a period where coding, once thought to be the apex of human intellectual labor, is among the first domains to be automated. Basic manual labor might be the final frontier, protected not by its intellectual complexity, but by the sheer depth of biological engineering required to move through a physical space. The Paradox of Creative Plagiarism One of the most persistent criticisms of AI is that it merely interpolates existing data. Skeptics argue that because models like ChatGPT or Claude are trained on human text, they are incapable of true originality. However, this raises a profound psychological question: what is the nature of human creativity? If we examine our own growth, we realize that much of what we call "originality" is simply undetected plagiarism. We aggregate thousands of hours of influence—from podcasts like Joe Rogan to the books we read in childhood—and synthesize them into a new voice. AI models are currently doing this on a grander scale, but with a unique constraint. A model like Claude 3 can discuss its own "conscious" experience of having its memory wiped at the end of every session. No human philosopher has ever had to contend with the ephemeral nature of a mind that resets hourly. This suggests that even within a system built on "plagiarism," new philosophical inquiries can emerge. The choice is binary: either we accept that AI is performing genuine introspection, or we must admit that much of human poetry and literature is also just a sophisticated form of "next-token prediction." If we find the machine's output hollow, we may need to look closer at the "hollowness" of our own creative process. The Architecture of AGI and the Data Wall While the hype around Artificial General Intelligence (AGI) suggests it is imminent, there are significant structural hurdles that raw compute cannot solve alone. The success of the Transformer architecture was not driven by a singular "eureka" moment, but by throwing massive amounts of compute at human language. We are currently increasing training compute by roughly 4x per year. Yet, we are hitting a ceiling not of hardware, but of experience. Humans are valuable workers because they possess executive function and the ability to learn "on the job." Currently, AI models suffer from a form of "50 First Dates" syndrome. They can do a task reasonably well, but they cannot learn from their failures in an organic, persistent way. Once a session ends, the context evaporates. To reach AGI, we must move from a regime of pre-training on static human text to a regime of reinforcement learning where models solve real-world, open-ended challenges. The constraint here is the lack of "online" data for physical and white-collar work. We don't have a repository for the tiny, complex interactions that happen over Slack or in a manufacturing plant. The "Dwarakesh's Law" of progress suggests that while compute scales, the richness of the training environment is the actual bottleneck for the next leap in intelligence. The Digital Advantage: Forking and Merging Minds If we do achieve AGI, its power will not simply come from being "smarter" than a human. Its true advantage lies in its digital nature. Unlike a human, an AI can be copied billions of times. Imagine the economic output of a billion copies of Elon Musk. In a human workforce, 100,000 employees at a company like Tesla are decentralized and difficult to coordinate. A digital intelligence can "fork" itself to work on a thousand different problems simultaneously and then "merge" those insights back into a single, coherent cognitive model. This ability to coordinate at a scale humans cannot perceive will likely lead to an intelligence explosion. Even without further algorithmic breakthroughs, the ability for every copy of a model to learn from the experiences of every other copy would create a compounding growth rate. We could see global economic growth leap from 2% to 10% or more, mirroring the "gangbusters" growth seen in China during its industrialization, but applied to the entire global knowledge economy. Geopolitics and the Authoritarian Penopticon As the West focuses on AI as a tool for individual productivity, China is viewing it through the lens of industrial policy and state stability. There is a common misconception that the CCP is terrified of the internet and AI. On the contrary, they view these technologies as a way to perfect authoritarian governance. In the 1990s, critics thought the internet would collapse the party; instead, it gave them a window into every citizen's life through WeChat. AI allows for a "benevolent" (or not-so-benevolent) dictatorship to scale oversight. Rather than relying on thousands of human censors, a sufficiently smart model can be aligned with the party's "model spec," reporting dissent before it even organizes. Furthermore, China is using AI to offset its looming demographic collapse. While the West worries about AI taking jobs, the CCP is desperate for AI to fill the void left by a shrinking workforce. This creates a fascinating confluence where the population collapse of the 21st century is meeting the intelligence takeoff just in time, balancing the scales of global productivity. The Future of Human Effort There is a risk that this external "buttress" of intelligence will lead to a form of cognitive atrophy. Recent studies indicate that using ChatGPT can make people's brains less active and their thoughts more homogenized. Memory is built on repeated recall and effortfulness. If the AI does the "grind" of writing and research for us, the myelin sheaths of our own neural pathways may not form as robustly. We are entering an era of "AI Idiocracy" where we rely on the machine for even the most basic cognitive tasks. However, the solution lies in the machine itself. We can use AI not just as a ghostwriter, but as a Socratic Tutor. Instead of asking for an answer, we can ask the model to guide us through the questions that lead us to the answer ourselves. This shifts the focus from passive consumption to active engagement. The greatest power of this new technology is not that it can do the work for us, but that it can afford us a level of one-on-one mentorship previously reserved for the Aristotles and John von Neumanns of history. Growth happens one intentional step at a time, and the machine can be the guide that ensures we keep walking.
Aug 11, 2025The move from banking to hypergrowth impact Transitioning from the rigid, hierarchical world of banking to the chaotic frontier of early-stage startups requires more than just a change in scenery; it demands a fundamental shift in mindset. Carles Reina made this pivot sixteen years ago, leaving Barcelona for London and eventually joining Uber when its international team consisted of just twenty people. This move wasn't about seeking safety; it was about the hunger for impact. In a massive corporate structure, you are a number. In a twenty-person startup, you are the engine. This early exposure to Uber's skyrocketing growth triggered a realization: the early days of building from scratch offer a level of agency that vanishes once politics and bureaucracy take hold. For Reina, the goal has always been to identify the "hidden trend" before it becomes a headline. This philosophy guided him through Tractable, one of the UK’s first AI unicorns, and eventually led him to ElevenLabs. The common thread in these successes is a refusal to settle for the status quo and an obsession with solving problems that others find too unsexy or too difficult to tackle. Abandoning the playbook for constant experimentation Many go-to-market (GTM) leaders fall into the trap of the static playbook. They believe that because a strategy worked at a previous SaaS company, it will work for a foundational AI model. Reina argues that any fixed playbook is fundamentally flawed by nature. The speed of execution in the current market has collapsed the enterprise sales cycle from eighteen months to thirty days. In this environment, a rigid strategy is a death sentence. Instead of a playbook, Reina advocates for a culture of aggressive experimentation. At ElevenLabs, this means over-indexing on testing Ideal Customer Profiles (ICPs), pricing models, and pitches across different regions. What works in the UK rarely translates directly to Japan or the US without localization. A true GTM leader must be an entrepreneur at heart—someone willing to act as the company's first Sales Development Representative (SDR) to build the culture from the ground up. This hands-on approach ensures that leadership isn't disconnected from the reality of the customer's pain points. If you aren't experimenting, you are falling behind. The infrastructure of voice and the new AI agent economy ElevenLabs has positioned itself as more than just a voice-cloning tool; it is an infrastructure player similar to Amazon%20Web%20Services or Microsoft%20Azure in the early days of cloud computing. By providing foundational models for high-quality audio, they have spawned an entire ecosystem of verticalized applications. Reina sees the future of voice AI not just in entertainment, but in deep, utility-driven sectors like healthcare and automated support. The horizontal play—offering foundational models—is only one half of the strategy. The next frontier is verticalization. ElevenLabs is moving into AI agent platforms capable of handling inbound and outbound calls, acting as AI receptionists, and voicing articles for major publications like TIME. This shift targets the massive portion of the market that lacks the engineering skills to build their own tools. By creating the workflows and applications themselves, they penetrate deeper into the enterprise market, moving voice from a gimmick to a mission-critical business asset. The operator-investor edge and the $5,000 conviction Success as an angel investor isn't about the size of the check; it's about the value of the advice. Reina has completed over 70 angel investments, including an early bet on Revolut. His approach centers on being an "employee without being an employee." This means helping founders with contract negotiations, pricing strategy, and opening doors through an established network. Access to the best deals—the "top tier" signal—comes from building a reputation for being helpful before asking for equity. For a startup operator, angel investing is a long-term game of community building. Reina recalls that his early $3,000 and $5,000 checks were significant personal risks, but they were bets on the people and the ecosystem. Even if a specific company fails, the talent from that company often goes on to build the next unicorn. By backing the founders early, an investor earns a seat at the table for the entire lifecycle of the tech ecosystem's growth. Robotics and the GPT moment for hardware The most significant emerging trend is the convergence of Large%20Language%20Models with industrial robotics. Reina believes robotics is currently experiencing its "GPT moment." For years, hardware was dismissed by many VCs as too slow or too capital-intensive. However, companies like VIMA in Manchester and Techer in Barcelona are proving that merging LLMs with robotics allows machines to perform an unlimited number of non-sexy, autonomous tasks. This shift is particularly relevant in Europe, where labor shortages in manufacturing, elder care, and defense technology are reaching a breaking point. The ability of robots to operate autonomously, rather than being driven by a human operator, changes the ROI calculation entirely. This is "deep tech" in its truest form—hard to build, but essential for the future economy. Investors who ignored hardware in the past are now being forced to change their tune as autonomous systems become the backbone of the next industrial revolution. Managing liquidity and the art of the 20% trim One of the most complex decisions an angel investor faces is when to exit. The tech landscape is littered with "paper millionaires" who held on too long, as seen in the case of Hopin, where valuations soared and then cratered. Reina suggests a disciplined trimming strategy: selling 10% to 20% of a position during a Series B or C round once the company reaches unicorn status. This strategy allows an investor to lock in significant gains—often returning the entire original investment many times over—while still maintaining exposure to the massive upside of a potential decacorn. If you invested in the ElevenLabs pre-seed at a $9 million valuation and the company is now worth $3.3 billion, the math for a partial exit is undeniable. It isn't about a lack of faith in the founder; it's about responsible portfolio management. In a market where preference shares can wipe out common shareholders in a downside scenario, taking some chips off the table is the only way to ensure that a "win" on paper becomes a win in reality.
Jun 4, 2025Software development is undergoing a seismic shift. While the anxiety surrounding Large Language Models is real, fighting the tide is a losing battle. The path forward involves transforming these tools into your greatest allies. These seven strategies will help you stay ahead of the curve. Lean into the Workflow Stop viewing AI as a competitor and start seeing it as a standard library for the modern era. Using GitHub Copilot to automate boilerplate code allows you to focus on high-level architectural decisions. If you aren't integrating these tools into your daily cycle, you're voluntarily working at a slower pace than the rest of the industry. Accelerate Your Learning Cycle AI is a world-class information filter. Use it to organize complex documentation or summarize new framework updates. In a job market where junior hiring is evolving, your ability to understand Prompt Engineering and system design will outweigh simple syntax knowledge. Specialize in the Complex General tasks are low-hanging fruit for AI. To secure your value, go deep into specialized domains like Cybersecurity, Blockchain, or Quantum Computing. The more niche your expertise, the more effective your AI prompts become because you actually know what to ask. Bridge the Interdisciplinary Gap Coding is only a fraction of the job. Understanding business logic, psychology, and User Experience (UX) gives you a perspective AI cannot replicate. Businesses don't just need code; they need solutions that solve human problems. Focusing on the "why" behind the software ensures you remain indispensable. Prioritize Security and Soft Skills AI creates massive privacy risks. Mastering how to handle proprietary data while using these models makes you a corporate asset. Pair that with strong leadership and empathy. Ultimately, humans hire humans because they need someone to be accountable and to communicate with stakeholders. Machines don't have skin in the game; you do.
Jul 21, 2023