Skip to content huecki Skip to content

huecki

Software, AI agents, messy notes and the occasional useful idea.

EXPERIMENTS

Current experiments

Small, focused projects where I test new workflows, interfaces, and infrastructure for AI agents in practice.

EXPERIMENT phase-flow replay + evidence honesty

https://

agent-buildprint.com

Agents no longer start from a vague assignment. They bootstrap a selected-buildprint packet, read the phase-flow constitution, write schema-valid runtime evidence, and cannot sell blockers as success.

$ agb start
→ phase before code
→ evidence before trust
→ replay before done
PHASE-FLOWEVIDENCEREVIEWSREPLAY
Open experiment →
EXPERIMENT

https://

media4agents.com

A dedicated project for media workflows built for AI agents.

MEDIAAGENTSWORKFLOWS
Open experiment →
AI ARCHITECT

https://

ai-architect.huecki.com

An evidence-led sparring partner for complex AI systems. Describe your case and get a concrete architecture with assumptions, trade-offs, failure modes, and traceable sources.

RAGAGENT RUNTIMESMEMORYEVALUATION
Open AI Architect →

AI Native Engineering

From prompt writer to AI system builder.

A self-paced learning path for developers who want to operate AI features, not just demo them — covering context budgets, Task Contracts, decomposition, evals, and fallbacks.

01

Tokens & Attention

Context windows, position effects, and lost-in-the-middle as real architecture constraints.

02

Context Engineering

Task Contracts, schemas, and source boundaries instead of longer prompts.

03

Agentic Delivery

Evals, traces, tool gates, and incident playbooks for operable AI features.

68 slides · self-paced · interactive

Becoming LLM-Native

Open the full learning path with the interactive slide deck, context models, Task Contracts, and operable AI-feature patterns.

Open AI Native Engineering →
· AI Agent Workflows

Your Agent Harness Needs a Release Process

A practical field note on operating agent-harness changes like product releases: start from a trace-backed failure, change one bounded component, evaluate repeated trials and private holdouts, then promote through review with a rollback path.

Read article →
· AI Agent Security

Coding Agents Need Hardened Harness Evals

Permissive coding-agent benchmarks hide a boring production truth: security policy changes agent behavior. Small teams should run the same task suite under nested hardening levels and separate model failures from tasks the policy made impossible.

Read article →
· AI Agent Workflows

Your Agent Harness Needs a Behavior Map

Harness Handbook points at a practical bottleneck in agent engineering: the behavior you want to change is scattered across prompts, state managers, tool calls, policy code, and tests. Build a behavior map before editing the harness.

Read article →
· AI Agent Workflows

Your Agent Eval Is Too Short

A final pass/fail score hides the part of agent work that matters most: where the run started drifting, whether it noticed, and whether it recovered. The practical replacement is a trajectory eval with checkpoints, failure labels, and recovery metrics.

Read article →
· AI Agent Security

Audit Local LLM Agents Like Runtimes

Local LLM agents can touch shells, files, browsers, credentials, memory, and messaging tools. Treat their runtime layer as source code worth auditing, then turn static findings into a manual review queue instead of automatic verdicts.

Read article →
· AI Agent Workflows

Your Agent Memory Test Is Probably Measuring the Wrong Thing

Most memory evals ask whether the agent got the final answer right. MemTrace suggests a sharper unit: one durable user fact tested across age, current state, earlier state, trajectory, and contradictory evidence. That turns memory from a vague feature into a small regression suite.

Read article →
· AI-first Engineering

Your Agent's Harness Is a Binary Now

Two 2026 papers from the same research lineage quietly retire prompt engineering as a discipline. The agent's system prompt is now a binary you can version, diff, and evolve with a 200-line loop. The four metrics that actually matter are not the ones your dashboard shows.

Read article →
· AI-first Engineering

Put an AI Slop Gate After Tests and Lint

Tests tell you whether behavior still works. Linters tell you whether code is syntactically and stylistically acceptable. An AI-slop gate catches the residue coding agents leave behind: fake comments, swallowed errors, any-casts, duplicated helpers, TODO stubs, and dead code.

Read article →
· AI-first Engineering

Your AI-Built UI Needs a Playtester, Not a Screenshot Review

AI-generated interfaces often look finished before they behave correctly. A GUI playtester loop uses a separate browser agent to interact with the artifact, record screenshots and action logs, turn broken flows into reproducible bug reports, and rerun the same script after repairs.

Read article →
· AI-first Engineering

Stop Judging AI Code by the Diff

Better AI coding is not mainly about better prompts. It is about the harness around the model: explicit contracts, separate builder and reviewer roles, evidence requirements, and a loop that turns failures into better specifications.

Read article →
· AI Agent Workflows

Give Your Agent Seatbelts, Not a Longer Prompt

When an agent keeps jumping from planning to editing to testing at the wrong time, the fix is not usually another paragraph of system prompt. Put the workflow into explicit states, give each state a tiny tool policy, and make phase changes visible.

Read article →
· AI Agent Workflows

AI Agents Need Evidence Before They Click

When an agent clicks, sends, pays, deletes, or extracts data, the critical truth cannot live only in model prose. Put a small evidence gate before risky tool calls: predicate, evidence type, source, decision.

Read article →
· AI Agent Workflows

Stop Asking AI to Critically Self-Check

Open-ended instructions like “critically self-check this” accidentally reward the model for producing criticism. The fix is not less review. It is calibrated review: explicit criteria, PASS_NO_CHANGE, evidence per finding, severity thresholds, and a tiny change budget.

Read article →
· AI-first Engineering

Your Onboarding Is Why Your Team Is Vibe Coding

Teams do not usually start vibe coding because developers became careless. They start because onboarding is broken: docs are stale, harnesses are undocumented, system knowledge lives in people’s heads, and AI turns missing context into plausible code and Markdown.

Read article →
· AI-first Engineering

The LLM-native developer needs more than prompts

The next developer skill is not writing clever prompts. It is building the operating system around LLMs: data quality, model versioning, evals, guardrails, incident response, review UX, and repo instructions agents can actually follow.

Read article →
· AI-first Engineering

Prompting Is Dead. Context Wins.

In 2026, good prompting is not about one magic sentence. The better approach is to curate context, define tools and schemas, set agent rules, and verify behavior with evals.

Read article →