kmail.at
← writing

ai

How to Become an AI Expert: The Path That Actually Produces One

Anthropic’s own postings say the quiet part: formal certifications are explicitly optional, and the screening question asks whether you have shipped an LLM system to production for real users. The full path: what the market actually hires for, the seven-layer builder’s stack, the learning science on why watching fails, the honest failure numbers — and a build ladder to start this month.

2026-10-05 · 21 min read

A craftsman’s workbench where a seven-layer stack of glowing modules is assembled by hand with precision tools and measurement instruments, dark scene with teal and green accents

How to Become an AI Expert: The Path That Actually Produces One

One of the frontier labs wrote the answer down. In April 2025, Anthropic published a job posting for a research engineer whose requirements include, verbatim, a section titled "Strong candidates need not have" — and its first line is "Formal certifications or education credentials." The same lab’s applied AI engineer postings screen every applicant with one question: "Have you personally built and shipped an LLM-powered application, agent, or workflows to production for real users (beyond personal experiments or proofs of concept)?" Not coursework. Not a certificate wall. Shipped systems, in production, for real users.

"Have you personally built and shipped an LLM-powered application, agent, or workflows to production for real users — beyond personal experiments or proofs of concept?" — the screening question on Anthropic’s applied AI engineer applications, 2026

That is what the market for AI expertise looks like in 2026, and it is nothing like what the course-selling corner of the internet describes. This article is the full map of the path that actually produces the experts employers compete for: what the hiring evidence says the title means, the seven-layer stack you have to be able to build end to end, the learning science on why watching does not work, the honest caveats — including the failure statistics that get misquoted everywhere — and a build ladder you can start this month.


What "AI expert" means when the market says it

The role is young enough that we know who named it. On June 30, 2023, swyx published "The Rise of the AI Engineer" on Latent Space, predicting that the engineers "working on productionizing AI APIs and OSS models" would "professionalize and converge on a title — the AI Engineer," and making a call that aged well: "This will likely be the highest-demand engineering job of the decade." His supply argument is worth quoting because it explains why the role is not machine learning: "There are ~5000 LLM researchers in the world, but ~50m software engineers." Andrej Karpathy’s version, pulled into that same essay: "One can be quite successful in this role without ever training anything." The honest footnote: swyx went on to found the AI Engineer conference, so read him as an interested party — but an interested party whose framing the hiring data has since caught up with.

The demand data is not subtle. LinkedIn’s Jobs on the Rise ranked Artificial Intelligence Engineer the number one fastest-growing job in the United States in 2025 and again in 2026 — the only AI role to top the list two years running, and the most common AI-engineering role globally. The detail worth more than the ranking is how the skills column moved between the two years: from "Large Language Models (LLM), Natural Language Processing (NLP), PyTorch" in 2025 to "LangChain, Retrieval-Augmented Generation (RAG), PyTorch" in 2026. The market re-priced the role from model-adjacent vocabulary to the builder’s stack in twelve months. Median prior experience for the role: about 3.6 to 3.7 years, with the top transitioned-from roles being full stack engineer, data scientist, and software engineer — not researchers retraining. The caveat that belongs in the same sentence: this is LinkedIn’s own platform metric, growth in title activity, not labor statistics; fastest-growing is not largest.

The breadth of demand corroborates it. Indeed’s Hiring Lab reported in January 2026 that postings mentioning AI finished 2025 at 134 percent above February 2020 levels while total postings finished just 6 percent above baseline — and in tech specifically, AI-mentioning postings were 45 percent above pre-pandemic levels while total tech postings were 34 percent below. Lightcast’s posting data (the dataset behind the Stanford AI Index’s labor chapter) counted specific generative-AI skill mentions jump from 16,000 in 2023 to 66,000 in 2024, with large-language-modeling mentions going from 5,000 to 20,000, and — for the first time since tracking began in 2010 — "artificial intelligence" as a skill cluster surpassing machine learning. Both are commercial platforms counting keyword mentions, not filled jobs; treat them as direction, not census.

Even the educators now define the title as a builder’s stack. Andrew Ng’s team published an AI Engineering Skills Map in August 2026 built explicitly to help developers prioritize what to learn and employers hire, and his terminology note is the right frame: "All developers — full-stack engineers, data engineers, DevOps engineers, machine learning engineers, and, yes, AI engineers — will need AI engineering skills," by analogy with how everyone needs cloud skills without holding a "cloud engineer" title. For the skills themselves, his summary is nearly a syllabus: understanding "LLMs, context engineering, RAG, agentic workflows" plus "how to use statistical techniques to measure, steer, and govern AI systems," where "a core skill is knowing how to drive disciplined evals and error analysis loops." Disclose the interest — DeepLearning.AI sells courses — and still note that this is the first-party educator landing on the same list the hiring postings use.

Then there is the survey that tells you what the job actually feels like. Stack Overflow’s 2025 Developer Survey found 84 percent of developers using or planning to use AI tools (up from 76 percent), 51 percent of professional developers using them daily — and in the same report, the distrust line: "More developers actively distrust the accuracy of AI tools (46%) than trust it (33%)." The single biggest frustration, cited by 66 percent: "AI solutions that are almost right, but not quite." Seventy-five percent say they ask another human when they don’t trust the AI’s answer. This is the market gap in three numbers: nearly everyone runs the tools, most people can’t fully trust the output, and the person who can systematically close that gap owns the title.

Which is precisely what the frontier labs interview for. OpenAI’s Applied AI Engineer posting for the Codex Core Agent team lists as core duties: "Work closely with research to develop and run evals to measure agent performance, regressions, failure modes, and edge cases" and "Improve performance through prompting, tool-use strategies, context construction, and model-facing experimentation." Anthropic’s applied posting names the same stack in its requirements: "Experience building LLM-powered tools or applications: prompting, context engineering, agent architectures, evaluation frameworks." Three labs’ postings compress into one sentence: nobody is hiring certificate holders; they are hiring people whose production failures have taught them things the rest of us haven’t learned yet.


The stack you have to be able to build: seven layers, end to end

Forget "learn Python for AI" listicles. The working definition this article uses: an AI expert is someone who can carry an idea across every layer between a raw model and a system that survives contact with production — and can explain what broke at which layer when it doesn’t. Six of the seven layers have unusually good primary-source anchors, published by the teams doing the work.

Why now, one paragraph: METR’s March 2025 study measured AI capability in the length of tasks agents can complete autonomously with 50 percent reliability, and found that metric "doubling approximately every 7 months for the last 6 years." METR’s own banner flags that the doubling figure reflects the data as originally published (their Time Horizons page carries the updated chart), so cite the trend, not today’s number. The anchors from that same paper: current models had "almost 100% success rate on tasks taking humans less than 4 minutes, but succeed <10% of the time on tasks taking more than around 4 hours." Each stack layer exists to stretch that horizon — and the people stretching it fastest are the ones learning the layers.

Layer 1 — The harness: workflows before agents

Anthropic’s "Building effective agents" (December 2024) is still the cleanest articulation of the first skill, and it is a discipline of restraint. Their taxonomy: "Workflows are systems where LLMs and tools are orchestrated through predefined code paths," while "Agents, on the other hand, are systems where LLMs dynamically direct their own processes and tool usage." Their recommendation cuts against every demo video: "finding the simplest solution possible, and only increasing complexity when needed. This might mean not building agentic systems at all." Behind the restraint sits real engineering depth — their rule of thumb is to "think about how much effort goes into human-computer interfaces (HCI), and plan to invest just as much effort in creating good agent-computer interfaces (ACI)." If you can design tool schemas that steer a model as deliberately as a good UI steers a human, you have the layer.

Layer 2 — The context: a budget, not a buffer

Anthropic’s September 2025 follow-up, "Effective context engineering for AI agents," names the resource you are actually managing: "Context, therefore, must be treated as a finite resource with diminishing marginal returns... LLMs have an ’attention budget’ that they draw on when parsing large volumes of context." And the definition to memorize: "good context engineering means finding the smallest possible set of high-signal tokens that maximize the likelihood of some desired outcome." Every failure report in the Stack Overflow survey — the "almost right" outputs — is largely a context problem wearing a model costume.

Layer 3 — The memory: treat the window like RAM

The MemGPT paper (Berkeley, October 2023) borrowed the operating system’s answer to exactly this problem: "we treat context windows as a constrained memory resource, and design a memory hierarchy [sic] for LLMs analogous to memory tiers used in traditional OSes." Its vocabulary is now industry furniture — main context (RAM) versus external context (disk), with the agent explicitly paging data in: "This out-of-context data must always be explicitly moved into main context in order for it to be passed to the LLM processor during inference." The Letta team’s one-liner about what MemGPT became — "the orchestration layer that sits above the model layer" — is a fair description of the whole skill.

Layer 4 — The retrieval: graphs for global questions

Microsoft’s GraphRAG paper (April 2024) states the ceiling of vanilla RAG in its first sentence: "RAG fails on global questions directed at an entire text corpus, such as ’What are the main themes in the dataset?’" — because that is summarization, not retrieval. Its fix is an index-time investment: "derive an entity knowledge graph from the source documents, then pregenerate community summaries for all groups of closely related entities," then answer via map-reduce over the summaries. The verified claim scope matters: "substantial improvements over a conventional RAG baseline" measured on "comprehensiveness and diversity" for "a class of global sensemaking questions" over million-token-range datasets — for local factual questions, plain RAG remains the right tool. Knowing which case you are in is the expertise.

Layer 5 — The evals: the layer that separates experts from tourists

Hamel Husain’s "Your AI Product Needs Evals" (his site; commonly dated April 2024) is the layer’s thesis statement, from someone who has watched enough teams fail to see the pattern: "unsuccessful products almost always share a common root cause: a failure to create robust evaluation systems." The working method is unfashionable and effective: "You must remove all friction from the process of looking at data," because "evaluation systems create a flywheel that allows you to iterate very quickly. It’s almost always where people get stuck when building AI products." And the layer has real judgment calls, not just tooling — on LLM-as-judge: track "the correlation between model-based and human evaluation," and when classes are imbalanced, "measure precision and recall separately" because raw agreement misleads. Notice that evals appeared verbatim in both frontier-lab postings above. Notice that no bootcamp teaches them.

Layer 6 — The model: know when you don’t need one

Honest disclosure: this is the layer with no dedicated canonical source in this article’s verified set — and that itself is instructive. The verified Karpathy line stands: one can be quite successful in this role without ever training anything. The actual expert skill at this layer is judgment about scope: when a hosted API, an open-weights local model (as with the decision models from our September piece), or a fine-tune is the right spend — and when the fix lives at another layer entirely. Karpathy’s training recipe advice applies to the whole stack, not just neural nets: "What we try to prevent very hard is the introduction of a lot of ’unverified’ complexity at once."

Layer 7 — The deployment: where sandboxing meets consequences

Anthropic’s agents post carries the deployment doctrine in two sentences: "We recommend extensive testing in sandboxed environments, along with the appropriate guardrails" — and METR’s conclusion supplies the stakes that make it a discipline rather than a checklist: capability trends pointing toward autonomously executed month-long projects "would come with enormous stakes, both in terms of potential benefits and potential risks." A person who has never watched their own agent misfire against real integrations does not yet have this layer.


Why watching doesn’t work: what the learning science says

The uncomfortable finding of expertise research is that the feeling of learning from consumption is mostly the confidence, not the skill. Ericsson, Krampe, and Tesch-Römer’s 1993 study of expert performers — the paper behind the pop "10,000-hour rule," which it never actually states — separates real practice from its imitation: "mere repetition of an activity will not automatically lead to improvement in, especially, accuracy of performance." What works is deliberate practice: "highly structured activity, the explicit goal of which is to improve performance. Specific tasks are invented to overcome weaknesses, and performance is carefully monitored to provide cues for ways to improve it further." Note the structural rhyme with evals: tasks invented to overcome weaknesses, performance carefully monitored. Building an eval loop for your own AI system is deliberate practice in the exact sense the paper means.

The education literature agrees at scale. Freeman et al.’s 2014 PNAS meta-analysis of 225 STEM studies found failure rates of 33.8 percent under traditional lecturing versus 21.8 percent under active learning — students in lecture courses were 1.5 times more likely to fail — with exam performance up 0.47 standard deviations under active learning. The paper is about classrooms, not tutorials, but passive-versus-active is the axis, and passive loses measurably. The testing-effect literature sharpens it: Roediger and Karpicke’s 2006 experiments found that "testing is a powerful means of improving learning, not just assessing it" — in the studied comparison, the group tested on the material recalled about 61 percent a week later versus roughly 40 percent for the group that re-read it more, despite the re-readers feeling more confident. Feeling productive while watching a tutorial is that confidence curve, with your name on it.

The AI field’s most popular teacher structures around this finding, knowingly. Karpathy’s advice page — written for students, apply everything here — says it plainly: "Reading and understanding IS NOT the same as replicating the content. Even I often make this mistake still: You read a formula/derivation/proof in the book and it makes perfect sense. Now close the book and try to write it down." His Neural Networks: Zero to Hero course is built as joint construction, per its own README: "we code and train neural networks together," with exercises attached to every lecture. His 2016 backprop essay names the trap for our stack specifically: "The problem with Backpropagation is that it is a leaky abstraction" — abstraction you haven’t rebuilt yourself will eventually leak, and then you will want the person who did. And his page closes the case for the portfolio over the transcript: "Getting actual, real-world experience, working on real code base, projects or problems outside of silly course exercises is extremely imporant [sic]... Document it well. Blog about it. These are the things people will care about a few years down the road."

The project-based-learning meta-analysis gives the build-path a number — Chen and Yang 2019, across 30 studies and 12,585 students, found a mean effect size of d+ = 0.71 for project-based learning against traditional instruction — with the honesty counterweight attached: other aggregations of build-style learning land lower (Visible Learning’s pooled mean across adjacent meta-analyses is ≈ 0.53), and large trials where implementation was weak came in near zero. Building beats watching on average; badly-run building does not beat anything. The Feynman anchor predates all of it and survives in his own handwriting on the Caltech blackboard photographed after his death: "What I cannot create, I do not understand." The four-step "Feynman technique" sold online is a listicle; the blackboard line and his "notebook of things I don’t know about" are the documented originals, and both say the same thing: creation is the exam.

The community slang for the failure mode is "tutorial hell" — that is folklore vocabulary, nobody famous’s quote, but it names a real measured phenomenon: consumption produces familiarity, replication produces capability, and only one of them shows up in an Anthropic screening question.


The honest caveats

An article about expertise that hid the failure data would be exactly the hype it warns against. So, with scope attached to every number.

The 95 percent statistic everyone quotes wrong. The claim circulating as "95% of AI pilots fail" comes from MIT Project NANDA’s "The GenAI Divide: State of AI in Business 2025" (v0.1, July 2025), whose actual wording: "Despite $30–40 billion in enterprise investment into GenAI, this report uncovers a surprising result in that 95% of organizations are getting zero return." The scope in the same breath: the report self-labels itself "Preliminary Findings," built from ~300 public initiatives, 52 organization interviews, and 153 conference-collected survey responses over six months — and the lead author told Fortune the 95 percent is of "the dataset," not of all companies. The critiques are documented, not invented: Futuriom called the figure presentation "irresponsible and unfounded" for lacking supporting data, and Wharton’s Kevin Werbach: "I’ve read through the document multiple times, and I still can’t understand where it comes from... If MIT Project NANDA stands behind the claims, it should release the full supporting data." The finding nobody disputes is more useful anyway: externally purchased tools reached production roughly twice as often as internally built ones. Read the 95 percent as directional evidence for a learning-and-integration gap — which is a market for experts, not a case against AI.

The agentic cancellation forecast, quoted fully. Gartner, June 2025: "Over 40% of agentic AI projects will be canceled by the end of 2027, due to escalating costs, unclear business value or inadequate risk controls." It is a vendor forecast resting partly on a webinar poll, and it names the failure mechanism precisely: "agent washing — the rebranding of existing products... without substantial agentic capabilities," with Gartner estimating "only about 130 of the thousands of agentic AI vendors are real." The same release forecasts 15 percent of day-to-day work decisions made autonomously by 2028 and agentic AI in a third of enterprise software. Cancellation and explosion are the same forecast: the field is being laundered, and the people who can tell real agentic capability from washed chatbots are the audit function.

The certificate question, stated without invented statistics. The verified market facts: the prompt-engineering hype cycle already collapsed — the Wall Street Journal reported that Indeed considers prompt-engineer postings "minimal" after 2023’s search surge, with per-million searches peaking near 144 in April 2023 and settling at 20–30 (secondary-verified; the WSJ piece is paywalled). Selection science (Sackett et al., 2022) re-based the validity estimates of hiring procedures and put structured interviews and job-specific evidence on top — though it studied classic selection procedures, not certificates, so treat the bridge as inference. And the verified absence: no rigorous public study establishes that AI course certificates predict on-the-job competence, in either direction. The burden of proof sits with the seller; the frontier labs put theirs on shipped artifacts.

The book that predicted the cycle, with its balance intact. Narayanan and Kapoor’s "AI Snake Oil" (Princeton, 2024) supplies the vocabulary: "AI snake oil is AI that does not and cannot work as advertised," and the media mechanism: "Many articles are just reworded press releases laundered as news." On generative AI specifically — and this is the balance the hype-weary usually drop — the authors grant the progress is real while judging the product: "immature, unreliable, and prone to misuse." Co-author Kapoor on the record elsewhere: "It’s hard to overstate the impact LLMs might have in the next few decades." Expertise is neither the panic nor the hype; it is the measured position in the middle, held by people who have read the failure reports and shipped anyway.


A build ladder you can start this month

Seven rungs, each one a portfolio artifact, ordered the way the market screens:

  • Rung 1 — Own one eval loop end to end. Pick a task you actually need done. Build the trace-review workflow, an error taxonomy, and a small judge you’ve validated against your own labels (precision and recall, not raw agreement — Husain’s checklist). This single artifact demonstrates the layer 80 percent of "AI engineers" don’t have.
  • Rung 2 — Ship one context-engineered pipeline. Take Anthropic’s "smallest possible set of high-signal tokens" as a design budget: instrument what enters the window per turn, cut one category of low-signal tokens at a time, and measure the delta on your eval. Publish the before/after.
  • Rung 3 — Add memory: implement the MemGPT pattern. Working context versus external context, explicit paging. Even a toy implementation — notes on disk, recited back into the window — teaches the constrained-resource reflex the papers describe.
  • Rung 4 — Build retrieval for questions that matter. Assemble a small graph index (entities, community summaries) over one corpus where global questions are real, and benchmark it against the vector baseline on both global and local queries. The comparison — including where plain RAG wins — is the artifact.
  • Rung 5 — Build one agent loop, workflow-first. Tools, environmental feedback, escalation thresholds; the Anthropic taxonomy as implementation. Ship the guardrail version where no verdict means no execution.
  • Rung 6 — Deploy locally and measure. Run an open-weights model end to end: latency, cost, failure modes, your own honest benchmark table. The measurement is the expertise; the hardware is incidental.
  • Rung 7 — Do it in public. Each rung gets a writeup: what broke, what the evals said, what you’d do differently. Karpathy’s advice page again: "Document it well. Blog about it. These are the things people will care about a few years down the road." That public trail is what the screening question reads.

One practitioner observation from my own builds, consistent with everything above: across the agent systems I’ve shipped and broken — including the decision-layer integration with a no-verdict-no-execution guard I described in September — the model was almost never the bottleneck. The context was, and the evals were. Every rung that looks like plumbing is where the leverage lived.


What expertise looks like from the inside

The expert is the person whose systems tell them when they are wrong. Everyone else’s systems just answer.

Strip the title to its mechanics and that is what remains. The eval loop, the attention budget, the memory tiers, the escalation threshold — these are all the same shape: instrumentation that converts silent failures into visible, priced ones. The learning science says the only reliable access runs through creation under feedback; the hiring data says the market pays for exactly that; the failure statistics say the demand is the verification, not the generation.

The METR trend — task horizons doubling every seven months, per their March 2025 publication with the live chart carrying updates — is usually read as a countdown over the role. Read it the other way: as the horizon stretches, the seven layers get more valuable, not less, because every extra hour of autonomous capability is an hour someone must be able to specify, bound, and verify. Nobody certifies their way into that. They build their way in, in public, one shipped failure at a time — and the screening question was ready before you finished reading: Have you personally built and shipped it to production for real users? The path is the answer.


Go deeper

Every claim above traces to a primary source, verified 2026-10-05. The market layer:

  • The Rise of the AI Engineer (swyx, Latent Space) — https://www.latent.space/p/ai-engineer
  • Andrew Ng: The AI Engineering Skills Map (The Batch #366) — https://www.deeplearning.ai/the-batch/issue-366/
  • Stack Overflow Developer Survey 2025, AI section — https://survey.stackoverflow.co/2025/ai/
  • LinkedIn Jobs on the Rise 2025 — https://www.linkedin.com/pulse/linkedin-jobs-rise-2025-25-fastest-growing-us-linkedin-news-gryie
  • LinkedIn Jobs on the Rise 2026 — https://www.linkedin.com/pulse/linkedin-jobs-rise-2026-25-fastest-growing-roles-us-linkedin-news-dlb1c
  • Indeed Hiring Lab, January 2026 — https://www.hiringlab.org/2026/01/22/january-labor-market-update-jobs-mentioning-ai-are-growing-amid-broader-hiring-weakness/
  • Lightcast: The Generative AI Job Market — https://www.lightcast.io/resources/blog/the-generative-ai-job-market-2025-data-insights
  • Anthropic: Research Engineer, ML (Reinforcement Learning) posting — https://job-boards.greenhouse.io/anthropic/jobs/4613568008
  • Anthropic: Applied AI Engineer, Beneficial Deployments posting — https://job-boards.greenhouse.io/anthropic/jobs/5413642008
  • OpenAI: Applied AI Engineer, Codex Core Agent posting — https://openai.com/careers/applied-ai-engineer-codex-core-agent-san-francisco/

The stack layers:

  • Anthropic: Building effective agents — https://www.anthropic.com/research/building-effective-agents
  • Anthropic: Effective context engineering for AI agents — https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
  • MemGPT: Towards LLMs as Operating Systems — https://arxiv.org/abs/2310.08560
  • Microsoft: From Local to Global (GraphRAG) — https://arxiv.org/abs/2404.16130
  • Hamel Husain: Your AI Product Needs Evals — https://hamel.dev/blog/posts/evals/
  • METR: Measuring AI Ability to Complete Long Tasks — https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/
  • METR: Time Horizons (live chart) — https://metr.org/time-horizons/

The learning science:

  • Ericsson, Krampe & Tesch-Römer 1993: The Role of Deliberate Practice — https://graphics8.nytimes.com/images/blogs/freakonomics/pdf/DeliberatePractice(PsychologicalReview).pdf
  • Freeman et al. 2014, PNAS: Active learning increases student performance — https://www.pnas.org/doi/abs/10.1073/pnas.1319030111
  • Roediger & Karpicke 2006: Test-Enhanced Learning — https://journals.sagepub.com/doi/10.1111/j.1467-9280.2006.01693.x
  • Chen & Yang 2019: Project-based learning meta-analysis — https://doi.org/10.1016/j.edurev.2018.11.001
  • Andrej Karpathy: advice page — https://cs.stanford.edu/people/karpathy/advice.html
  • Karpathy: A Recipe for Training Neural Networks — https://karpathy.github.io/2019/04/25/recipe/
  • Karpathy: Yes you should understand backprop — https://karpathy.medium.com/yes-you-should-understand-backprop-e2f06eab496b
  • Karpathy: Neural Networks Zero to Hero — https://github.com/karpathy/nn-zero-to-hero

The honest caveats:

  • MIT Project NANDA: The GenAI Divide, State of AI in Business 2025 (v0.1) — https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf
  • Futuriom: the critique of the 95% figure — https://futuriom.com/articles/news/why-we-dont-believe-mit-nandas-werid-ai-study/2025/08
  • Gartner: agentic AI cancellation forecast — https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027
  • Narayanan & Kapoor: AI Snake Oil (official excerpt, Princeton UP) — https://pup-assets.imgix.net/onix/images/9780691249643/9780691249131.pdf
  • Sackett et al. 2022: revisiting selection-validity estimates — https://doi.org/10.1037/apl0000994

Companion from this blog:

  • Laya vs Jev: Open Weights and the Sovereignty Question (September) — https://kmail.at/blog/laya-vs-jev-open-weights-sovereignty
  • Context Engineering: The Discipline That Ate Prompt Engineering — https://kmail.at/blog/context-engineering-the-discipline-that-ate-prompt-engineering

← all writing

keep reading