*A grounded-answering thin slice, from React to FastAPI and OpenAI.*
> **One focused capability, engineered end to end:** a user provides Context and a Question; the application returns either a grounded answer with literal supporting Evidence or an explicit Abstention when the Context is insufficient.
This project is an early but complete milestone in my progression into **AI application engineering**. It extends my existing full-stack and solution-engineering foundation into Python, FastAPI, typed model integration, AI trust boundaries, and systematic verification.
It is intentionally a small educational system—not a production product. The value is in the engineering fundamentals it makes visible and testable.
## At a glance
| | |
| ------------------- | ----------------------------------------------------------------- |
| **Status** | Completed against its agreed thin-slice scope |
| **Core capability** | Grounded Answer with 1–3 literal Evidence excerpts, or Abstention |
| **Stack** | React, TypeScript, FastAPI, Pydantic, OpenAI Responses API |
| **Architecture** | Typed frontend → API → service → replaceable provider |
| **Automated proof** | 60 backend tests + 5 frontend tests |
| **Live proof** | Real browser-to-OpenAI smoke test and recorded demonstration |
| **Delivery** | 13 of 13 planned milestones completed |
| **Source code** | Public on GitHub — [juanjoarranz/grounded-answering-thin-slice](https://github.com/juanjoarranz/grounded-answering-thin-slice) |
## Why I built it
Calling a model is easy. Building a dependable application around a model requires more:
- What inputs are accepted?
- What output states are legal?
- When must the system refuse to answer?
- How is model output checked before it reaches the user?
- How can the provider be replaced during testing?
- How do failures remain useful without exposing internal details?
I postponed retrieval, embeddings, and orchestration frameworks so I could first understand and implement these foundations directly. The result is a small system whose complete request path can be explained, tested, and challenged.
## The user experience
The interface accepts a **Context**, a **Question**, and an optional language tag. The Context is the only permitted knowledge source.
The application produces one of two successful domain outcomes:
1. **Grounded Answer** — a non-empty answer plus one to three unique Evidence excerpts copied exactly from the Context.
2. **Abstention** — no partial answer and no Evidence when the Context is missing, incomplete, or too contradictory.
The UI also handles loading, prevents duplicate submissions, and displays controlled errors. An Abstention is treated as responsible system behavior, not as a technical failure.
## Architecture: guarded from input to output
![[image.png]]
*The request crosses explicit validation gates before a typed result returns to the React interface.*
The end-to-end flow is deliberately shallow:
1. React sends a typed request to `POST /api/answer`.
2. FastAPI and Pydantic normalize and constrain untrusted input.
3. `AnswerService` delegates asynchronously through a small `AnswerProvider` protocol.
4. The OpenAI adapter requests a structured `AnswerResponse` from the Responses API.
5. Pydantic enforces the only two valid response states: Grounded Answer or Abstention.
6. The service verifies that every Evidence item is an exact substring of the original Context.
7. FastAPI returns the typed result or maps the failure to a controlled HTTP response.
### Engineering choices that matter
**Contracts before prompts.** Request limits, response invariants, and error semantics are defined in code rather than left to prompt wording.
**The Context is untrusted data.** Instructions found inside it are not treated as system commands. Context and Question remain separated from the system instruction boundary.
**Structured output is necessary, but not sufficient.** Schema-valid output can still be unsupported. The application therefore adds a domain check that Evidence must occur literally in the supplied Context.
**The provider is replaceable.** Business orchestration depends on a minimal protocol rather than directly on the OpenAI SDK. Tests inject a fake provider; production composition resolves the real adapter.
**Failures are part of the public contract.** Invalid input, timeouts, unavailable configuration, and invalid provider output map to controlled responses without leaking secrets, prompts, full Context, raw model responses, or SDK internals.
## Verification: proof at multiple boundaries
![[image-1.png]]
*Fast isolated checks cover application behavior; a final smoke test proves the real integration path.*
| Verification layer | Result | What it demonstrates |
| --- | --- | --- |
| Backend | **60 tests passed** | Schemas, service rules, API behavior, settings, provider adapter, dependency composition, and error mapping |
| Frontend | **5 tests passed** | Form accessibility, loading, duplicate prevention, answer rendering, Abstention, and controlled errors |
| Python quality | **Black + Ruff passed** | Formatting and static quality gates |
| Frontend quality | **Oxlint + TypeScript + Vite build passed** | Static analysis, type checking, compilation, and production bundling |
| Real integration | **Grounded Answer + Abstention passed** | Browser, CORS, HTTP, FastAPI, real OpenAI call, and React rendering |
The automated suites make no real OpenAI call: the provider and browser boundaries are replaced so the checks remain fast, deterministic, and free. The manual smoke test then exercises the one path those test doubles cannot prove.
## Three-minute demonstration
![[FastAPI Thin Slice demo.mp4]]
The walkthrough shows local startup, an answerable question with literal Evidence, an unanswerable question that triggers Abstention, and the main architectural boundaries.
## What this project demonstrates
- **AI application engineering:** integrating a model inside an application contract rather than exposing a raw prompt call.
- **Full-stack delivery:** connecting a React/TypeScript interface to an asynchronous Python/FastAPI backend.
- **Schema and API design:** using Pydantic to constrain input and enforce cross-field response invariants.
- **Trust-aware design:** separating instructions from untrusted Context and validating Evidence against the original source.
- **Testable architecture:** isolating external providers behind a small seam and replacing boundaries with deterministic fakes.
- **Operational judgment:** distinguishing successful Abstention from failure and translating expected errors without leaking internals.
- **Scope discipline:** proving one capability thoroughly before adding retrieval, frameworks, persistence, or infrastructure.
## Honest scope and limitations
This project does **not** claim production readiness. It deliberately excludes RAG, document ingestion, authentication, persistence, streaming, automatic retries, deployment, monitoring, load testing, and a second model-based verifier.
Literal Evidence is a useful guardrail, but it is not formal proof that every answer claim is entailed by the cited text. The real-provider path is also manual, and the initial frontend suite is intentionally small.
These are visible boundaries, not hidden gaps. Recording them is part of the engineering work and defines the next iteration.
## Next evolution
The next slice will preserve the current contracts and provider seam while adding:
1. document ingestion and retrieval-backed Context;
2. a versioned evaluation set for answerable, partial, contradictory, unanswerable, and prompt-injection cases;
3. runtime validation of frontend responses;
4. deployment, observability, cost, latency, resilience, and security evidence.
The progression is deliberate: **first make one AI request understandable and trustworthy; then expand the system without losing those properties.**
## Professional direction
My transition into AI Engineering builds on established experience with Microsoft 365, SharePoint, React, .NET, and solution delivery. This thin slice is public evidence of that evolution: learning Python and modern model APIs while retaining the habits that make software maintainable—clear boundaries, explicit contracts, automated tests, controlled failures, and honest technical trade-offs.
For more, visit my [professional profile](https://cv.juanjoarranz.info) or my [GitHub profile](https://github.com/juanjoarranz).
<span class="cta">[View source on GitHub](https://github.com/juanjoarranz/grounded-answering-thin-slice) [[Portfolio|Back to case studies]]</span>