An essay
Structure Is All You Need?
Helping AI agents construct programs in unfamiliar languages

It would be incredibly difficult to deny the capabilities of the current generation of large language models and their impact on programming and software development as a whole.
A familiar advantage
AI has, without a shadow of a doubt, changed how we write programs forever. In fact, I would go so far as to argue that it may also influence which languages we choose to program in. We know that AI learns from the data that is fed into it during training. The availability of data for a specific task can affect how well a model performs at that task. Want a model to get good at C++? Well, show it a tonne of C++ code. Want clean AI-generated Erlang gen_servers? Well, give it some Erlang, and you'd expect it to get better.
Essentially, the performance of a model on a certain language can depend in part on how much data from that language was consumed during training. This may help explain why AI is better at generating code for extremely popular languages like C++, Rust, TypeScript, and Python, but not as good at generating code for languages like, say, EYG, without first looking up the documentation.
The adoption problem
As AI-assisted coding becomes more common, a language's "AI readiness" may influence its adoption. Popular languages may have an advantage here, with established communities and more code examples for models to learn from. Newer programming languages with fewer users and examples could find it harder to offer the same level of AI assistance.
As a result, newer programming languages may face a steeper hill to climb in the era of AI-assisted coding. Early adopters might need to write more code manually or spend more time guiding their AI tools while the language's examples and tooling develop. That could make a new language harder to sell to users who already rely on AI assistance in a familiar language, though its other benefits may still make the switch worthwhile.
That being said, I am far from convinced that we have invented all the programming languages that will ever be written. There is always room for innovation in this domain, and new programming languages have the potential to change how we think about writing programs. The next Rust, Zig, Java, or Haskell is out there waiting to be invented, and it might have something useful to offer to us as programmers, even in the era of AI-assisted programming.
If this becomes a barrier, what can we do about it? How could we level the competition for newer languages on the market and foster some level of AI readiness for said languages?
Just the rules
One idea I have been tinkering with is to move the grammar rules of the language into the tooling the AI agent uses and relieve the model of the burden of remembering its syntax. A grammar defines how to write a syntactically correct program in that language. In this implementation, those rules are encoded in a renderer that turns the model's description of a program into source code.
For example, here are simplified grammar subsets of a handful of programming languages and code they can generate:
Python: see full grammar
function ::= "def" NAME "(" parameters ")" ":" NEWLINE
INDENT "return" expression NEWLINE DEDENT
parameters ::= NAME { "," NAME }
expression ::= atom { "+" atom }
atom ::= NAME | INT
def add(x, y):
return x + y
Haskell: see full grammar
declaration ::= [ signature NEWLINE ] equation NEWLINE
signature ::= NAME "::" type
type ::= "Int" | "Int" "->" type
equation ::= NAME NAME { NAME } "=" expression
expression ::= atom { "+" atom }
atom ::= NAME | INT
add :: Int -> Int -> Int
add x y = x + y
Gleam: see full grammar
function ::= [ "pub" ] "fn" NAME "(" parameters ")"
"->" "Int" "{" expression "}"
parameters ::= parameter { "," parameter }
parameter ::= NAME ":" "Int"
expression ::= atom { "+" atom }
atom ::= NAME | INT
pub fn add(x: Int, y: Int) -> Int {
x + y
}
A language's grammar rules constrain the programs that can be generated from it. Whenever we write a program, we invoke our learned knowledge of those rules to express our ideas in code. For newer languages with very few samples to train on, that familiarity can be harder to acquire.
But what if the model doesn't need to remember the language's syntax rules? What if, for a new language, we could devise a way to generate source from the model's description of the program alone?
To test this out, I built a programming language of my own and tried to see if I could get AI to generate a valid program in it.
The programming language I built is called Benzene and aims to be a statically typed functional programming language powered by a Hindley-Milner (HM) type-inference engine. The type checker infers types and checks their consistency before runtime.
To qualify as valid in the context of the language's formalism, a program has to have the following two properties:
- It has to be syntactically correct.
- It has to satisfy the type checker.
Below is a sample of Benzene's grammar definition and a couple of valid programs one can generate from it.
Benzene: see full grammar
program ::= { function }
function ::= "func" NAME "(" [ parameters ] ")" ":>" "Int"
{ binding | expression } "end"
parameters ::= parameter { "," parameter }
parameter ::= NAME ":" "Int"
binding ::= "let" NAME [ ":" "Int" ] "=" expression
expression ::= atom { "+" atom }
atom ::= call | NAME | INT
call ::= NAME "(" [ arguments ] ")"
arguments ::= (NAME | INT) { "," (NAME | INT) }
Here, NAME represents an identifier and INT an unsigned integer literal.
This illustrative subset covers the examples below; the full grammar supports
additional declarations, types, and expressions.
func add(x: Int, y: Int) :> Int
x + y
end
func add(x: Int, y: Int) :> Int
x + y
end
func main() :> Int
let total: Int = add(20, 22)
total
end
To my knowledge, Benzene has not been used outside its own development, so models may have had little exposure to it. I cannot verify whether its development materials appeared in training data, but its limited use makes it a useful test case for generating code in a language a model may be unfamiliar with.
At the heart of this effort is ether, Benzene's compiler. The compiler enforces the grammar rules during lexing and parsing. After an untyped abstract syntax tree (AST) has been generated, several modules such as scope resolution, symbol resolution, and type checking are applied in series over the data structure to produce a fully typed and resolved AST.
This language cannot yet run programs, as IR generation is still in development. However, for the purpose of this exercise, passing the compiler's parsing, resolution, and type checks is a satisfactory checkpoint to validate the current idea.
Let the tool handle the syntax
Within the language's tooling, I have developed an MCP server whose purpose is to relieve a model of the burden of assembling Benzene syntax and give it access to the compiler for validation. The implementation is written in Gleam, and you can take a look at the MCP folder here.
Essentially, instead of asking the model to write a function by remembering all the syntax rules of Benzene, I ask it to describe the structure of that function. What is its name? What parameters does it take? What operations should its body perform? These are still decisions the model has to make. The tool takes those decisions and turns them into source code.
The server exposes two tools: generate_program for constructing source from structured inputs, and check_program for
checking source that has already been written. When an agent discovers the tools through tools/list, it receives a JSON
Schema describing the available constructs and their fields. For example, here are the arguments it could pass to
generate_program to describe the addition function from earlier:
{
"construct": {
"kind": "Function",
"identifier": "add",
"param_list": [
{ "param": "x", "data_type": { "kind": "NamedType", "name": "Int" } },
{ "param": "y", "data_type": { "kind": "NamedType", "name": "Int" } }
],
"return_type": { "kind": "NamedType", "name": "Int" },
"body": [
{
"kind": "Binary",
"operator": "Add",
"left": { "kind": "Identifier", "name": "x" },
"right": { "kind": "Identifier", "name": "y" }
}
]
}
}
Notice that the model has not supplied func, :>, or end. It has described a function, its parameters, and an addition.
The server decodes this JSON into structured Gleam values, checks that their formation is allowed, and then renders them.
The entry point for generation is quite small:
pub fn generate(
construct: ct.Construct,
) -> Result(String, List(GenerationError)) {
case validate(construct) {
[] -> Ok(ct.to_string(construct, 0))
errors -> Error(errors)
}
}
The formation checks catch things like invalid names, duplicate parameters, declarations in the wrong position, and case branches with the wrong number of patterns. If these checks fail, the tool returns errors with paths into the construct so the agent can see where it went wrong. If they pass, the renderer supplies the concrete syntax. The request above becomes:
func add(x: Int, y: Int) :> Int
x + y
end
This becomes more useful when the syntax has rules that are easy to miss. For one, Benzene does not support parentheses
for grouping expressions. If I want to add x and y before multiplying the result by two, the renderer needs to preserve
that order. Here is the corresponding construct using the same Gleam values that JSON requests are decoded into:
ct.Binary(
ct.Multiply,
ct.Binary(ct.Add, ct.Identifier("x"), ct.Identifier("y")),
ct.Integer(2),
)
That expression renders as:
{ x + y } * 2
Without the braces, x + y * 2 would describe a different calculation. The renderer uses a helper called primary to wrap
nested operations in scoped expressions. The same helper handles call arguments, where Benzene requires primary expressions.
This is the actual call-rendering branch:
Call(name, arguments) ->
name
<> "("
<> {
arguments
|> list.map(fn(arg) { primary(arg, depth) })
|> string.join(", ")
}
<> ")"
For example, ct.Call("add", [ct.Binary(ct.Add, ct.Integer(1), ct.Integer(2)), ct.Integer(3)]) renders as
add({ 1 + 2 }, 3). The model chooses what it wants to calculate, while the renderer handles how that calculation must be
written in Benzene. It also handles string escaping, comment escaping, and the punctuation of functions, types, and collections.
The syntax rules are encoded in this renderer; the server does not read an EBNF document and automatically derive a generator.
Of course, producing the right syntax does not mean the model has chosen the right names or types. The agent can still ask
for a function that does not exist, or try to add a string to an integer. This is where the compiler comes back into the picture.
By default, generate_program sends the generated source to ether scan through stdin. The compiler lexes, parses, resolves
names, and checks types, then returns diagnostics and inferred symbol types to the tool.
The response includes the source, a validated flag saying whether a compiler report was obtained, and a separate valid
flag saying whether that report contains no errors. A failed check gives the agent feedback it can use to revise its request.
For individual fragments, validate: false skips the compiler check but still performs formation checks.
By moving syntax construction into the tool, the model no longer needs to remember how to write Benzene source. The details of writing the source have been moved out of the model: grouping, escaping, delimiters, and other syntax conventions. The tool constructs its syntax, and the compiler checks the result.
An example
Here is a recording of the process end to end. In this example, the agent uses the MCP server to generate Benzene source and obtain compiler feedback. This is the construction and validation loop described above, applied to an actual request.
Watch or download the recording.
The agent uses build_lexer.py to assemble the JSON request; the Gleam/Erlang MCP server generates and validates the Benzene source.
In this recording, I prompted the agent to write a tokenizer for the Benzene programming language itself.
Limitations
One could argue that the problem has only moved: the model still needs to understand how to describe its intended program. However, it now does so in JSON, a format likely to be more familiar to LLMs. This moves the problem into a domain where errors may be easier to diagnose, while still letting us generate useful programs in our newly created language.
To be fair, many syntax errors are fairly benign, and clear compiler diagnostics can make them easy to fix.
This approach might be more practical for languages with a more esoteric design, where unfamiliar syntax conventions
would otherwise need repeated guidance. In Benzene, for example, both x + y * 2 and { x + y } * 2 can pass the
compiler's checks while describing different calculations. The renderer preserves the structure the model described.
Even so, I can already see some limitations of this approach. One is token consumption. Below is a comparison of the tokens needed to describe a program in JSON and those needed to write its source directly.
How many tokens?
For a small comparison, I used gpt-6.1-sol in Codex to construct the following program and its JSON equivalent.
I sent the JSON through generate_program, and the server returned exactly the same source, with compiler validation passing.
func add(x: Int, y: Int) :> Int
x + y
end
func main() :> Int
let total: Int = add(20, 22)
total
end
| Representation | Tokens | Relative to source |
|---|---|---|
| Direct Benzene source | 45 | 1.00× |
| Compact JSON tool arguments | 190 | 4.22× |
| The same arguments, pretty-printed | 378 | 8.40× |
| Full compact JSON-RPC request | 217 | 4.82× |
Even without formatting whitespace, describing this program in JSON takes 145 more tokens than writing its source. So for this example, moving syntax construction into the tool comes with a larger text representation.
See the JSON
{"construct":{"kind":"Module","data":[{"kind":"Function","identifier":"add","param_list":[{"param":"x","data_type":{"kind":"NamedType","name":"Int"}},{"param":"y","data_type":{"kind":"NamedType","name":"Int"}}],"return_type":{"kind":"NamedType","name":"Int"},"body":[{"kind":"Binary","operator":"Add","left":{"kind":"Identifier","name":"x"},"right":{"kind":"Identifier","name":"y"}}]},{"kind":"Function","identifier":"main","return_type":{"kind":"NamedType","name":"Int"},"body":[{"kind":"Let","identifier":"total","data_type":{"kind":"NamedType","name":"Int"},"value":{"kind":"Call","identifier":"add","arguments":[{"kind":"Integer","value":20},{"kind":"Integer","value":22}]}},{"kind":"Identifier","name":"total"}]}]}}
I counted the text with tiktoken 0.14.0, using o200k_base. The installed tokenizer has no verified mapping
for this model, so these are counts under that encoding, rather than exact model billing (see
OpenAI's token-counting documentation). The compact JSON has no
formatting whitespace. The direct source also passed ether check. The JSON-RPC row includes the transport envelope, which an
MCP client can supply without the model writing it.
Well, how does the overhead scale?
To look at this beyond one example, I generated seven increasingly larger programs using the same family of calculations.
Each added stage introduces an integer-increment function and a binding in main that passes the previous result to it.
The sizes use 1, 2, 4, 8, 16, 32, and 64 stages. Every JSON request passed through the MCP server, produced exactly the
expected source, and passed compiler validation.
For this family, each stage adds 29 source tokens and 111 JSON tokens. The multiplier stays close to 3.83 times as the programs get larger.
However, that straight line is partly a result of how the experiment was constructed: I kept adding the same kind of structure. A different mix of expressions, types, or literals could change the ratio. The counts also do not measure whether the tool reduces the effort of writing a program.
See the counts for each size
| Stages | Source tokens | JSON tokens | JSON / source |
|---|---|---|---|
| 1 | 41 | 156 | 3.80× |
| 2 | 70 | 267 | 3.81× |
| 4 | 128 | 489 | 3.82× |
| 8 | 244 | 933 | 3.82× |
| 16 | 476 | 1,821 | 3.83× |
| 32 | 940 | 3,597 | 3.83× |
| 64 | 1,868 | 7,149 | 3.83× |
The code generation model used in this was gpt-6.1-sol. I used the same o200k_base encoding and compact tool-argument format
as above. The source also passed ether check directly.
Does the tool actually help?
Quite an obvious alternative is to give the model the grammar and a few examples, then let it write Benzene directly and use the compiler to repair mistakes. To test that, I ran a small comparison of the two approaches.
For this comparison, I gave gpt-6.1-sol, with low reasoning effort, eight tasks in two modes:
- A: write Benzene directly, then use
check_programfor feedback. - B: describe the program in JSON, then use
generate_programwith compiler validation.
Both modes received the full grammar and three examples. The tasks covered addition, grouped arithmetic, expressions in call arguments, Boolean operations, lists and tuples, string escaping, case branches, and anonymous functions. I ran each task once in each mode.
| Approach | First-attempt compiler passes | Task matches* | Repair rounds | Total tokens |
|---|---|---|---|---|
| A: direct source | 8 / 8 | 8 / 8 | 0 | 322,499 |
| B: structured generation | 8 / 8 | 8 / 8 | 0 | 350,169 |
*I reviewed the generated source against each task. Both approaches matched the requirements, including expression grouping and the exact escaped string.
Well, the direct approach worked just as well here. Both passed on the first attempt, so there were no repairs for the tool to save. Structured generation used about 8.6% more total tokens. For these small tasks, this experiment does not show a reliability advantage or token savings from using the generator.
Per-task counts and experiment setup
| Task | A: total tokens | B: total tokens |
|---|---|---|
| Addition | 40,014 | 43,447 |
| Grouped arithmetic | 41,584 | 44,662 |
| Expressions in call arguments | 40,031 | 43,608 |
| Boolean operations | 39,834 | 43,509 |
| Lists and tuples | 41,071 | 44,601 |
| String escaping | 38,275 | 41,590 |
| Case branches | 39,982 | 43,455 |
| Anonymous functions | 41,708 | 45,297 |
Both modes also received the input schema for their respective tool. Each task started with fresh context and allowed
up to three repair rounds. A harness submitted each model response to the actual MCP server and returned the complete
tool response for acknowledgement or repair. The model generated ordinary source or JSON without constrained decoding.
Each completion used a fresh, read-only Codex session with that trial's full history replayed. The experiment measures
generation and feedback, rather than autonomous tool discovery. All sixteen final sources also passed ether check directly.
These totals come from Codex's reported input and output usage, including prompts, schemas, tool feedback, and the model's final acknowledgement. Cached input is included, and reasoning tokens are included in output rather than added again; the reported reasoning count was zero in this run.
Codex's own instructions and the unabridged JSON-RPC feedback make up much of the input. A used 322,041 input tokens and 458 output tokens; B used 348,439 input tokens and 1,730 output tokens. Of those inputs, 200,704 tokens were cached for A and 202,752 for B. The total-token ratio therefore depends on this setup and should not be confused with the earlier JSON-to-source ratio.
Eight small tasks with one run each leave plenty untested. Since both approaches passed everything, this test cannot tell us where either starts to struggle. Larger programs, repeated runs, and a wider mix of constructs would be needed to see whether handling syntax in the renderer saves enough repair work to justify the overhead.
What if the model has to learn the language on the go?
For a second comparison, I started fresh agents with no Benzene grammar, examples, or tool schemas in their initial
context. They received the same eight tasks, written without Benzene syntax, and could choose when to discover the
tools and read the documentation. Both approaches had access to the same documents and the same three-repair limit.
I again used gpt-6.1-sol, with low reasoning effort, and one run per approach per task.
| Approach | First-attempt passes | Repair rounds | Documentation requests | Total tokens |
|---|---|---|---|---|
| A: direct source | 8 / 8 | 0 | 8 | 468,080 |
| B: structured generation | 8 / 8 | 0 | 8 | 501,541 |
Both approaches still passed every task on the first submission, and when I checked them by hand, all sixteen programs matched the requirements. The agents consulted documentation themselves; giving them no grammar upfront did not make these tasks difficult enough to reveal a reliability advantage. Structured generation used about 7.1% more total tokens in this setup.
What the agents read, and how the experiment ran
| Task | A: documents requested | B: documents requested | A: tokens | B: tokens |
|---|---|---|---|---|
| Addition | grammar | all | 57,736 | 62,882 |
| Grouped arithmetic | grammar | all | 59,113 | 64,582 |
| Expressions in call arguments | all | examples | 58,421 | 55,793 |
| Boolean operations | all | all | 58,185 | 63,468 |
| Lists and tuples | all | all | 59,464 | 64,547 |
| String escaping | all | all | 56,698 | 61,582 |
| Case branches | all | all | 58,372 | 63,420 |
| Anonymous functions | all | all | 60,091 | 65,267 |
The trials used an external agent loop: the model chose actions, and a harness executed documentation requests and program submissions against the real MCP server. No grammar or examples were inserted until the agent requested them. Each action used a fresh CLI completion with that trial's history replayed; there was no context from this essay or from another task. This tests learning and generation through the loop, rather than native MCP integration.
The totals are counted the same way as before and also include discovery and the requested documentation. A recorded 466,847 input and 1,233 output tokens; B recorded 499,098 input and 2,443 output tokens.
What if there is only compiler feedback?
Starting without documentation in context still allowed the agents to look it up. This time, I removed that
option altogether. Fresh agents received no grammar or examples and could not browse or read files. A could
only submit source to check_program and learn from compiler feedback. B could only use generate_program,
its structured input schema, and its formation and compiler feedback. Neither could request language documentation.
I kept the same eight tasks, gpt-6.1-sol with low reasoning effort, and an initial attempt plus three repairs.
| Approach | First-attempt passes | Final task passes* | Repair rounds | Total tokens |
|---|---|---|---|---|
| A: compiler feedback alone | 0 / 8 | 0 / 8 | 24 | 764,715 |
| B: generator schema and feedback | 5 / 8 | 8 / 8 | 3 | 406,410 |
*Passing means the compiler accepted the program and, when I checked it by hand, it matched the task.
In this case, the direct agents exhausted their available repair budgets without producing any compiler-accepted program, which is honestly expected. Compiler errors alone, especially without documentation, severely limit how many of the grammar rules the model is able to infer on its own. The structured agents, on the other hand, produced programs matching all eight tasks.
This is a different question from whether the tool beats writing source with documentation available. The generator's schema supplies structural information that compiler feedback alone does not. Under this restriction, the tool helped; in the two tests with documentation, direct generation already worked.
Per-task results and setup
| Task | A: result / repairs | B: result / repairs | A: tokens | B: tokens |
|---|---|---|---|---|
| Addition | Fail / 3 | Pass / 0 | 90,644 | 42,041 |
| Grouped arithmetic | Fail / 3 | Pass / 0 | 97,034 | 42,446 |
| Expressions in call arguments | Fail / 3 | Pass / 0 | 91,655 | 41,742 |
| Boolean operations | Fail / 3 | Pass / 1 | 92,274 | 64,422 |
| Lists and tuples | Fail / 3 | Pass / 1 | 102,324 | 67,952 |
| String escaping | Fail / 3 | Pass / 0 | 86,970 | 40,234 |
| Case branches | Fail / 3 | Pass / 1 | 94,734 | 63,634 |
| Anonymous functions | Fail / 3 | Pass / 0 | 109,080 | 43,939 |
The same external agent loop dispatched requests to the actual MCP server. Documentation was removed from tool
discovery and unavailable at the server, rather than merely discouraged. Initial prompts, every submitted candidate,
compiler report, and model-usage record were retained. All successful programs also passed standalone ether check.
The totals are counted the same way as before. They also include failed attempts, so they are bounded workflow costs
rather than costs for equally successful outputs.
So what then?
At the very least, we have been able to generate a valid program in a new programming language that satisfies our compiler's frontend requirements, albeit in a rather token-expensive manner. The nice thing here, though, is that in the absence of proper documentation, code generation is still possible via the MCP server.
In all fairness, though, the resulting data confirm what we already knew about these capable models: with proper documentation, you really don't have to worry much about the novelty of your problem. We've been able to show that, at least for these small tasks, just the documentation alone defeats the need for the tool itself and even incurs a lower token cost when compared to programs generated via the MCP.
The examples do show this is a possible solution. Is it currently impractical for larger programs from a token-economics standpoint? Yes. But possible nonetheless. Besides, I personally think this is a cool idea, and I had a fun time implementing it.
Additionally, this naive implementation can always be improved on, and even better ideas are never too far away.
I am still a firm believer that we are yet to discover all the useful programming languages and ideas that will ever exist. Even in the era of agentic software development, there's always room for innovation that newer languages can bring to the programming landscape.
For now, this is just traditional codegen exposed via MCP. But who knows what it could look like later down the line? Regardless, I am excited to see what newer programming languages will be invented in a future where humans may be doing little to no software implementation themselves.
Materials