Support Rotations Creep Into Engineering Velocity Supporting large-scale data infrastructure is a continuous battle against ambiguous priorities. At Pinterest, engineers faced an endless stream of complex distributed system failures. Troubleshooting Apache Spark requires navigating noisy logs, complex metrics, and scattered runbooks. Standard solutions rely on engineers manually correlation of trace data. Draško Proferović, a Staff Engineer at the company, realized that while human time is strictly finite, Large Language Model (LLM) capabilities can scale on demand. This spark of insight led to the creation of Medic for Apache Spark, an autonomous agent built to diagnose job failures. The Breakdown of the Single-Prompt Agent The initial prototype used a single reasoning-and-acting (ReAct) agent. It relied on a massive prompt containing solving guidelines, formatting requirements, and historical error examples. It fell flat. Prompt tuning became an unsustainable chore; adjusting instructions in one area instantly broke performance in another. Out-of-memory errors and giant trace logs quickly choked the LLM's context window. To fix this, the team integrated LangFuse for tracing and built an end-to-end test harness that recorded production state, transforming chaotic live errors into reproducible local fixtures. Compressing Raw Data into Token-Friendly Assets To prevent noisy logs and massive metrics from consuming the agent's context window, the engineers built two pipelines. First, they deployed an exception classifier that clusters exceptions and filters out benign red herrings. Second, they isolated metrics analysis into a dedicated sub-agent. Instead of passing raw time-series data, they converted the metrics into annotated dashboard images. The sub-agent reasons visually over these graphs, returning a highly compressed text summary to the coordinator. Orchestrating specialized nodes with LangGraph Today, Medic operates as a multi-agent cooperative system built on LangGraph. A triage agent classifies user intent and determines the job's lifecycle state. Parallel research agents investigate specific failure hypotheses, scoring their findings before passing them to a supervisor. A healer agent then pulls remediations directly from internal runbooks. This modular architecture allows the team to easily scale the system to support other engines like Flink and Trino.
Companies
Jan 2022 • 1 videos
High activity month for Pinterest. Chris Williamson among the most active voices, with 1 videos across 1 sources.
Apr 2024 • 1 videos
High activity month for Pinterest. 20VC with Harry Stebbings among the most active voices, with 1 videos across 1 sources.
Nov 2025 • 1 videos
High activity month for Pinterest. Chris Williamson among the most active voices, with 1 videos across 1 sources.
Jan 2026 • 1 videos
High activity month for Pinterest. Morning Brew Daily among the most active voices, with 1 videos across 1 sources.
Mar 2026 • 1 videos
High activity month for Pinterest. The Prof G Pod – Scott Galloway among the most active voices, with 1 videos across 1 sources.
Jun 2026 • 1 videos
High activity month for Pinterest. AI Engineer among the most active voices, with 1 videos across 1 sources.
Jul 2026 • 1 videos
High activity month for Pinterest. AI Engineer among the most active voices, with 1 videos across 1 sources.
- 5 days ago
- Jun 2, 2026
- Mar 16, 2026
- Jan 28, 2026
- Nov 1, 2025
The Conviction to Scale the Impossible OpenAI didn't emerge from a vacuum; it was born from a radical bet on two factors that much of the tech world initially dismissed: deep learning and the predictive power of scale. Sam%20Altman notes that while he was interested in AI since childhood, the actual conviction to launch the venture seven years ago came from seeing that bigger was consistently better. The industry was skeptical. Many viewed the project as a binary risk—it would either work spectacularly or fail completely. This skepticism didn't deter the founding team; it motivated them. They pursued an attack vector rooted in the belief that if they could keep doing things previously thought impossible, they were on the right track. Brad%20Lightcap, who joined as the company's first business-minded hire, saw a unique property in the research. Unlike other moonshots like nuclear fusion or quantum computing, OpenAI showed a trajectory of incremental, predictive improvement. This wasn't just a blind leap of faith. It was a data-driven pursuit of a technological revolution. Today, that revolution has manifested as the fastest-scaling company in history, reaching over $2 billion in revenue in a timeframe that has left traditional SaaS benchmarks in the dust. The Anatomy of a High-Octane Partnership The relationship between Sam%20Altman and Brad%20Lightcap provides a blueprint for leadership in high-growth environments. Altman, despite his role, identifies as a non-operator. He prefers the strategic, long-term orientation of an investor, focusing on the "one to three things" that act as the fastest accelerants to the future. His role is to maintain a maniacal focus on the horizon, ensuring the company doesn't lose its innovative edge as it scales. In contrast, Lightcap manages the "how." He stepped into the COO role with a willingness to build out entire business functions from scratch, even when no playbook existed for selling advanced AI to the enterprise. This partnership thrives on high-bandwidth communication and a clear division of labor. Altman handles the research-to-product vision, while Lightcap builds the market infrastructure. They move fast because they are aligned on the global bets, allowing Lightcap to make dozens of daily decisions independently without clogging the Altman bottleneck. This decentralized execution is what allows the organization to maintain velocity even as its complexity explodes. The Steamroller Problem: Startup Strategy in the Age of AGI For entrepreneurs and venture capitalists, the most pressing question is how to build in a world where OpenAI is constantly shipping updates that can wipe out entire product categories. Sam%20Altman is blunt about this: if you build assuming the current model (like GPT-4) is the ceiling, you will be steamrolled. Many startups focus on fixing the "little things" or building wrappers around current limitations. This is a losing strategy because OpenAI's mission is to solve those very limitations at the base layer. The winning strategy is to build assuming GPT-5, GPT-6, and beyond will continue on a steep trajectory of improvement. Successful founders ask themselves: "Would a 100x improvement in the underlying model make my product better or make it obsolete?" If your business benefits from the model becoming more intelligent, more personalized, and more deeply integrated into the user's life, you are safe. If your business depends on the model remaining "dumb" or limited in specific ways, you are in the path of the steamroller. The enduring value for startups will not be in the base model, which is rapidly becoming a commodity, but in the personalization and deep workflow integration that a general-purpose provider cannot replicate at scale. Solving the Compute and Intelligence Bottleneck The primary constraints on OpenAI's growth aren't market demand or competition; they are physical and scientific. To provide abundant, near-zero-cost intelligence to every person on Earth, the company requires a massive, coordinated effort across the entire hardware stack. This includes chips, data centers, and power. Altman views this as a "whole system problem." While the cost of intelligence is falling, the demand for it is scaling even faster. The goal is to drive the cost of high-quality intelligence so low that it transforms society. Currently, the models simply aren't smart enough to solve the world's most complex problems, such as curing cancer or accelerating scientific breakthroughs to a point where we view 2024 as "barbaric." The fix is one-dimensional: increase the underlying intelligence. This requires a relentless focus on research. Within the OpenAI culture, research drives product, and product drives sales. There is no compromise on this hierarchy. If the research fails to innovate, the business stops growing. Enterprise Adoption and the ROI Trap Brad%20Lightcap has observed a recurring mistake in how large corporations approach AI. Many enterprises attempt to force AI into existing business processes to achieve a quantifiable, line-item ROI—like cutting 20% of supply chain costs. While valuable, this approach misses the broader impact. The real return comes from the "supply of time" shift. When an employee who used to spend two days on a task now finishes in two minutes, it frees them for higher-order work. This impact is harder to quantify on a balance sheet but is transformative when scaled across 100,000 employees. Enterprises that treat the current models as static tools are setting themselves up for failure. They should instead view AI as a rapidly evolving platform. The organizations that will win are those that set up flexible workflows capable of absorbing the next wave of intelligence as soon as it drops. Adoption isn't a one-time event; it's a continuous integration of increasing intelligence into the corporate DNA. The Future of Growth and Talent Scaling at this speed requires a specific type of talent. While OpenAI is currently the "hottest" company in tech, Altman and Lightcap are wary of hiring mercenaries. They look for mission-oriented individuals who are determined, communicative, and capable of fast iteration. Interestingly, the company skews slightly older than the typical Silicon Valley startup, particularly in its research and leadership teams. This is a byproduct of the depth required to push the boundaries of science. Altman's growth mindset has evolved as well. He admits that ChatGPT's success broke many traditional rules of growth. When you are in the midst of a once-in-a-generation technological revolution, the standard retention curves and marketing playbooks become secondary to the utility of the product itself. The future of OpenAI is one of genuine abundance. Despite the geopolitical and socioeconomic instability Altman sees in the world, he remains bullish on the ability of AI to level the playing field, providing every individual with the tools to do amazing things. This isn't just a business for them; it's a mission to ensure AGI benefits all of humanity, shifting us from a world of scarcity to one of unlimited potential.
Apr 15, 2024The Internal Cost of Challenging Power True growth often requires stepping into the line of fire. When James O'Keefe discusses his experiences with Project Veritas, he isn't just talking about investigative reporting; he is describing a psychological battle against institutional inertia. The recent FBI raid on his home, triggered by the acquisition of Ashley Biden's diary, represents a profound psychological threshold. Facing a government battering ram at 6:00 AM creates a level of stress that would break most individuals. This is where resilience moves from a concept to a survival mechanism. Psychologically, the impact of federal litigation and raids cannot be overstated. O'Keefe openly admits to experiencing symptoms of PTSD after being shackled in New Orleans years ago. Yet, he views this suffering as a necessary precursor to impact. In the world of personal development, we often speak about the "comfort zone," but O'Keefe operates in the "conflict zone." He posits that if you are not being challenged by those in power, you are likely not fulfilling your purpose as a disruptor. This mindset shifts the perspective from being a victim of the system to being a catalyst for its transparency. The Paradox of Relative Deception One of the most complex psychological landscapes explored by O'Keefe is the ethics of undercover work. He introduces the "paradox of relative deception," a choice between deceiving a subject to reveal a hidden truth or deceiving an audience by withholding that truth. From a psychological standpoint, this is a classic moral dilemma. To reach a higher state of collective awareness, one must sometimes adopt a persona that is fundamentally untruthful. This mirrors the internal negotiations we all make. We often wear masks in our professional or personal lives to achieve specific outcomes. O'Keefe argues that the "clown world" of modern media—where CNN or The New York Times operate with inherent biases—forces a non-traditional approach. By recording subjects who believe they are in a private setting, he bypasses the social scripts and defensive mechanisms people use to protect their interests. This is psychological archaeology, digging beneath the surface of public relations to find the raw, unvarnished human motive. Overcoming the Silence of the Majority O'Keefe frequently references Aleksandr Solzhenitsyn to explain the prevailing atmosphere of fear in modern society. He suggests that we are living through a "tragedy of the commons" regarding truth. While 98% of people may recognize a lie, they are often paralyzed by the fear of losing their status, their social media accounts, or their livelihoods. This collective silence allows a small minority to dictate the narrative. Breaking this silence requires a specific type of courage. O'Keefe notes that Whistleblowers like Eric Cochran from Pinterest or Frances Haugen demonstrate a contagious form of bravery. When one person stands up, it creates a "domino effect" of integrity. This is the essence of mindset shifts: moving from a state of self-preservation to a state of principle-preservation. Most individuals are "surviving at any price," sacrificing their values to maintain their comfort. The path to true potential lies in the opposite direction—being willing to lose the temporary for the sake of the eternal truth. The Legal and Ethical Mirror Operating under constant scrutiny requires a radical level of self-awareness. O'Keefe’s internal rule for his staff is to behave as if a jury is always in the room. This is a powerful psychological tool for habit formation and ethical conduct. When we act under the assumption that our private lives will eventually become public, we naturally align our actions with our stated values. This eliminates the cognitive dissonance that plagues many in corporate or political environments. Despite dozens of lawsuits, O'Keefe highlights that Project Veritas has never lost a defamation case. He finds a strange sense of peace in depositions, viewing them as opportunities for absolute transparency. While the FBI and The New York Times seek to unearth his secrets, he claims his only secrets are the identities of his sources. This total lack of personal concealment acts as a shield. If there is no gap between who you are and what you do, your enemies have nothing to grip. This is a masterclass in living an integrated life, where the external pressure only serves to harden the internal resolve. Conclusion: The Future of Trust As public trust in mainstream institutions like Pfizer and the Department of Justice continues to erode, the demand for unvarnished truth will only grow. O'Keefe’s work, detailed in his book American Muckraker, suggests that the future of journalism—and perhaps society—depends on individuals who are willing to be the "boogeyman" to those in power. The goal is not merely to win a legal battle, but to win the battle for the human conscience. By refusing to be intimidated and continuing to publish the "truth unspoken," we can move toward a society grounded in reality rather than manufactured narratives. The path forward is through the fire, not around it.
Jan 29, 2022