Building real software with LLMs is fundamentally different from chatting with a bot. In production applications, you need **predictable schema outputs**, **low perceived latency through streaming**, and **strict guardrails** to prevent prompt injections.
1. Type-Safe Structured Outputs
Never rely on standard natural text parsing with regex when extracting data. Modern foundation models support native JSON schemas and strict function calling.
from pydantic import BaseModel
from openai import OpenAI
class CodeReviewResult(BaseModel):
summary: str
severity: str # 'low' | 'medium' | 'critical'
suggested_fix: str
performance_impact: bool
client = OpenAI()
completion = client.beta.chat.completions.parse(
model="gpt-4o-mini",
messages=[
{"role": "system", "content": "You are an expert code review engine."},
{"role": "user", "content": "Review this function: async def get_data(): ..."}
],
response_format=CodeReviewResult,
)
# result is 100% typed and guaranteed to match schema
review: CodeReviewResult = completion.choices[0].message.parsed
print(f"Severity: {review.severity}, Fix: {review.suggested_fix}")
2. Perceived Latency with Streaming (SSE)
LLMs generate responses token-by-token. Instead of making your users wait 4-8 seconds for a complete response, stream tokens directly to the browser via Server-Sent Events (SSE) or WebSockets.