Module 04.2: Representing Functions¶
Module 04.1 introduced one of the central ideas of this course:
Code is data.
We represented arithmetic expressions as recursive Haskell data, then wrote an eval function that recursively traversed those expressions and gave them meaning.
Now we take the next step.
Functions have been central to everything we have done so far. We have passed functions as arguments, returned them as results, partially applied them, composed them, and treated them as ordinary values.
So if code is data, and functions are code, then we should be able to represent functions as data too.
Functions are data.
This module extends our little expression language so that it can represent function definitions and function applications. To give those new expressions semantics, we will introduce substitution, use it to explain function application, compare call-by-name and call-by-value evaluation, and then reconnect all of this to currying and multi-argument functions.
By the end of this module, you should be able to:
- represent function definitions and applications in an abstract syntax tree,
- explain why a function definition is itself a value,
- describe function application as substitution followed by evaluation,
- read and perform expression substitution using notation such as
e[x → e'], - implement a recursive substitution function over our expression datatype,
- explain why
evalnow needs to return anExprather than only aDouble, - distinguish call-by-name from call-by-value,
- explain some tradeoffs between those evaluation strategies,
- represent multi-argument functions using nested one-argument functions,
- explain why nested application requires evaluating the function position first, and
- identify some important problems that our substitution-based evaluator still does not solve.
How to use this module
When you see a box labeled Gradescope question, answer that question in the Module 4.2 Completion assignment on Gradescope.
Other boxes labeled Pause and think, Try it, or Practice are there to help you understand the material. There is nothing to submit for those unless the box explicitly says Gradescope question.
Recap: Syntax, Abstract Syntax, and Semantics¶
Before adding functions, let's reconnect to the picture from Module 04.1.
A programmer begins with concrete syntax:
1 + 2 * 3
A parser turns that concrete syntax into abstract syntax, represented as a tree.
An evaluator then assigns meaning to that tree:
concrete syntax abstract syntax semantics
1 + 2 * 3 ---> + ---> 7
/ \
1 *
/ \
2 3
In CS 131, our current focus is mostly on the right-hand side of this pipeline:
abstract syntax ---> semantics
We use:
- Haskell data types to represent abstract syntax, and
- a Haskell
evalfunction to define the semantics.
Parsing, the concrete-syntax-to-abstract-syntax step, comes later.

Our arithmetic language so far¶
We represented arithmetic expressions with:
data Op
= PlusOp
| MinusOp
| TimesOp
| DivOp
deriving (Show, Eq)
data Exp
= Num Double
| BinOp Exp Op Exp
deriving (Show, Eq)
For example:
BinOp (Num 1) PlusOp
(BinOp (Num 2) TimesOp (Num 3))
is the abstract syntax for:
1 + 2 * 3
We can think of this as a tiny programming language, which we'll call:
Arith, short for Arithmetic Expressions.
And:
eval :: Exp -> Double
defined the semantics of Arith.


We then added variables and a lookup mechanism.
At that point, we could represent and evaluate:
- numeric expressions,
- arithmetic operations,
- variables.
What is still missing?
Two large pieces:
- turning source strings into ASTs, which is parsing, and
- functions.
Parsing comes later.
Functions are today's problem.


What Can We Do With a Function?¶
Before designing a datatype, ask a simpler question:
What are the fundamental things we do with functions?
For our purposes, there are two.
1. Define a function¶
When we define a function, we specify:
- a parameter, and
- a body describing the computation.
For example:
addFive x = x + 5
Here:
x parameter
x + 5 body
For now, we will assume every function has exactly one parameter.
That might sound restrictive, but recall currying from Module 02.2. We already know that a function that appears to take several arguments can be understood as a sequence of one-argument functions.
2. Apply a function¶
Once we have a function, we can apply it to an argument:
addFive 22
The argument does not have to be a literal value:
addFive (22 - 4)
It can be an arbitrary expression.
So the two structures our language needs to represent are:
function definition:
parameter + body
function application:
function + argument

Representing Functions as Data¶
Let's extend Exp.
type VarName = String
data Exp
= Var VarName
| Num Double
| BinOp Exp Op Exp
| FunDef VarName Exp
| Apply Exp Exp
deriving (Show, Eq)
We have added three important constructors.
Var¶
Var "x"
represents a reference to a variable named x.
FunDef¶
FunDef "x" body
represents a function whose:
- parameter is named
"x", and - body is
body.
For example:
FunDef "x"
(BinOp (Var "x") PlusOp (Num 1))
represents the function:
x ↦ x + 1
or, in informal Haskell-like notation:
\x -> x + 1
Apply¶
Apply functionExpression argumentExpression
represents applying one expression as a function to another expression as its argument.
For example:
Apply
(FunDef "x"
(BinOp (Var "x") PlusOp (Num 1)))
(Num 5)
represents applying our "plus one" function to 5.

(The image above spells this constructor FuncDef; this page consistently uses FunDef — same structure, just the shorter spelling.)
One subtle point: these functions are anonymous¶
Notice that:
FunDef "x" body
stores the parameter name, but not a name for the function itself.
There is no equivalent of:
addOne x = x + 1
where addOne is available inside or outside the body.
Instead, a FunDef is simply a value representing a function.
For now, that works well. Later, it will become important when we ask how a function could refer to itself recursively.
What Does It Mean to Apply a Function?¶
We have represented function application syntactically:
Apply
(FunDef "x"
(BinOp (Var "x") PlusOp (Num 1)))
(Num 5)
But syntax is only half the story.
We need semantics:
What should this expression mean?
Gradescope question: What should function application do?
Submit this response on Gradescope.
Describe in words what you think should happen if we evaluate this expression:
Apply
(FunDef "x"
(BinOp (Var "x") PlusOp (Num 1)))
(Num 5)
We know the result should eventually be:
Num 6
But how do we get there?
The basic idea is familiar from ordinary algebra.
We have:
function:
x ↦ x + 1
argument:
5
To apply the function:
- find occurrences of the parameter
xin the body, - replace them with the argument
5, - evaluate the resulting expression.
So:
x + 1
becomes:
5 + 1
which evaluates to:
6
In our abstract syntax:
BinOp (Var "x") PlusOp (Num 1)
becomes:
BinOp (Num 5) PlusOp (Num 1)
and then:
Num 6

This replacement operation is called substitution.
Expression Substitution¶
We need a precise way to say:
Replace every appropriate occurrence of variable
xin expressionewith expressione'.
The standard notation is:
e[x → e']
Read this as:
"In expression
e, substitute expressione'for variablex."
Some examples:
(x + 1)[x → 5]
= 5 + 1
(z * z + 7)[z → 2]
= 2 * 2 + 7
(x + y + z)[y → (x / 3)]
= x + (x / 3) + z
An extremely important point:
Substitution itself does not evaluate the expression.
For example:
(x + 1)[x → 5]
becomes:
5 + 1
during substitution.
It does not become 6 until evaluation happens afterward.
![A slide titled "Notation" defining expression substitution as e[x → e'] ("substitute the expression e' for variable x in the expression e"), with three worked examples — (x + 1)[x → 5] = 5 + 1, (z * z + 7)[z → 2] = 2 * 2 + 7, (x + y + z)[y → (x / 3)] = x + (x / 3) + z — and the note "No evaluation happens, just substitution."](../img/04.2/substitution-notation.png)
Practice Substitution Before Implementing It¶
Before writing a recursive function, work through the structural cases by hand.
Gradescope question: Perform these substitutions
Submit this response on Gradescope.
What do these four expressions become after performing the substitution?
A.
10[x → 5]
B.
x[x → 5]
C.
x[y → (1 + 2)]
D.
(x + (3 / x))[x → 12]
Think about why each behaves differently.
A number¶
10[x → 5] = 10
There are no variables inside the number 10, so there is nothing to replace.
The matching variable¶
x[x → 5] = 5
We found exactly the variable being replaced.
A different variable¶
x[y → (1 + 2)] = x
We are replacing y, not x, so this occurrence stays unchanged.
A compound expression¶
For:
(x + (3 / x))[x → 12]
we recurse into the expression's pieces:
(x[x → 12]) + ((3 / x)[x → 12])
and eventually obtain:
12 + (3 / 12)
Again, substitution stops there. Computing the numeric result is a separate evaluation step.
![A slide titled "Next: let's implement! Wait. Let's think first..." showing the four substitution exercises with their answers and case labels: 10[x → 5] = 10 (base case: nothing to substitute), x[x → 5] = 5 (base case: just the replacement expression), x[y → (1 + 2)] = x (base case: nothing to substitute), and (x + (3/x))[x → 12] = (x[x → 12]) + ((3/x)[x → 12]) = 12 + (3/12) (recurse: substitute on left and right).](../img/04.2/substitution-exercises-answers.png)
Implementing Substitution¶
Now the implementation follows the recursive structure of Exp.
We will write:
substitute replacement variable expression
so that:
e[v → e']
corresponds to:
substitute e' v e
Before looking at the complete function, try to fill in its pieces.
Gradescope question: Implement substitute
Submit this response on Gradescope.
How can you implement the substitute function?
Expression substitution:
e[v → e'] == substitute e' v e
Fill in the blanks:
substitute :: _____ -> _____ -> _____ -> _____
substitute e' (Var v) (Num n) =
_________
substitute e' (Var v) (Var w) =
________________________
substitute e' (Var v) (BinOp left op right) =
_____________________________________
substitute _ _ _ =
error "bad substitution"
The core implementation is:
substitute :: Exp -> Exp -> Exp -> Exp
substitute _ (Var v) (Num n) =
Num n
substitute replacement (Var v) (Var w) =
if v == w
then replacement
else Var w
substitute replacement var (BinOp left op right) =
BinOp
(substitute replacement var left)
op
(substitute replacement var right)
substitute _ _ _ =
error "bad substitution"
The first three cases correspond almost directly to our hand-worked examples:
number -> unchanged
matching Var -> replacement
different Var -> unchanged
BinOp -> recurse into both children

But our language contains more than numbers, variables, and binary operations¶
We have also added:
FunDef
Apply
We'll deliberately postpone these cases at first.
For nested function definitions, we will eventually need to recurse into the function body:
substitute replacement var (FunDef varname body) =
FunDef varname (substitute replacement var body)
This simple rule will let us make progress today.
However, it is not the final word on substitution.
Near the end of the module, we will see an example where blindly substituting inside every FunDef can do the wrong thing because function parameters introduce their own bindings.
That problem is intentional. It points directly toward the next major topic of the course.
Why Does Evaluation Now Return an Exp?¶
Back in Module 04.1, our evaluator had type:
eval :: Exp -> Double
That made sense when every fully evaluated expression was a number.
But now consider:
FunDef "x"
(BinOp (Var "x") PlusOp (Num 1))
What should evaluating a function definition produce?
It should produce... the function.
A function is already a value.
So Double can no longer describe every possible result of evaluation.
Instead:
eval :: Exp -> Exp
We stay inside the world of expressions.
Numbers evaluate to numeric values represented as Exp:
eval (Num x) = Num x
Function definitions evaluate to themselves:
eval (FunDef varname body) =
FunDef varname body
And binary operations evaluate their children and then combine the resulting numeric expressions:
eval (BinOp left op right) =
evalOp (eval left) op (eval right)
with:
evalOp :: Exp -> Op -> Exp -> Exp
evalOp (Num x) PlusOp (Num y) = Num (x + y)
evalOp (Num x) MinusOp (Num y) = Num (x - y)
evalOp (Num x) TimesOp (Num y) = Num (x * y)
evalOp (Num x) DivOp (Num y) = Num (x / y)
This explains why the result of our earlier example is:
Num 6
rather than merely:
6
We want evaluation to have one result type capable of representing all of our language's values, including functions.


Evaluating Function Application¶
Now we can finally add semantics for:
Apply f arg
The simplest version directly mirrors our informal description.
If the function position already contains a function definition:
eval (Apply (FunDef varname body) arg) =
eval (substitute arg (Var varname) body)
Read it in three stages:
1. take the function body
2. substitute the argument for the parameter
3. evaluate the resulting expression
For example:
Apply
(FunDef "x"
(BinOp (Var "x") PlusOp (Num 1)))
(Num 5)
becomes:
BinOp (Num 5) PlusOp (Num 1)
and then:
Num 6

But there is another perfectly reasonable choice.
When Should the Argument Be Evaluated?¶
Compare these two definitions.
Strategy 1¶
eval (Apply (FunDef varname body) arg) =
eval (substitute arg (Var varname) body)
The argument is substituted without being evaluated first.
Strategy 2¶
eval (Apply (FunDef varname body) arg) =
let val = eval arg
in eval (substitute val (Var varname) body)
The argument is evaluated first, and then the resulting value is substituted.
These are different evaluation strategies.
The first is call-by-name.
The second is call-by-value.

Gradescope question: Compare the two evaluation strategies
Submit this response on Gradescope.
We gave two possibilities for evaluating a function application:
eval (Apply (FunDef varname body) arg) =
eval (substitute arg (Var varname) body)
and:
eval (Apply (FunDef varname body) arg) =
let val = eval arg
in eval (substitute val (Var varname) body)
What is the difference between these two evaluation strategies? Will they always produce the same result? Are there cases where one is worse or better than the other?
Call-by-Name¶
Under call-by-name, the argument expression is substituted directly into the function body without first being evaluated.
Suppose the argument is:
someExpensiveComputation
and the parameter appears three times:
x + x + x
Substitution may produce:
someExpensiveComputation
+ someExpensiveComputation
+ someExpensiveComputation
If evaluation then computes each copy independently, we may repeat the same expensive work several times.
So call-by-name has an important potential cost:
It can duplicate unevaluated computation.

Call-by-Value¶
Under call-by-value, we first evaluate:
arg
to a value:
val
and substitute that value into the body.
This avoids recomputing the argument every time the parameter appears.
But there is another possible waste.
Consider a function whose body never uses its parameter:
x ↦ 10
If we apply it to a huge computation, call-by-value computes that argument before discovering that the result does not depend on it at all.
Even more dramatically, suppose the argument never terminates.
Call-by-value will get stuck evaluating the argument, even though the function body might not need it.
So call-by-value has a different potential cost:
It can evaluate an argument that never needed to be evaluated.


Connection back to laziness¶
This should sound familiar from Module 02.2.
Haskell's lazy evaluation shares an important motivation with call-by-name: do not evaluate an expression until its value is actually needed.
But ordinary Haskell also avoids naively recomputing the same demanded argument every time it is used. It uses sharing, often described as call-by-need.
You do not need to implement call-by-need here. The point is that a feature that earlier looked like a quirky fact about Haskell now appears as a deliberate answer to a programming-language design question:
When should function arguments be evaluated?
Functions with More Than One Argument¶
So far, our datatype assumes:
FunDef parameter body
with exactly one parameter.
What about:
add x y = x + y
?
We already know the answer from currying.
A two-argument function can be represented as a function that takes one argument and returns another function:
FunDef "x"
(FunDef "y"
(BinOp (Var "x") PlusOp (Var "y")))
Conceptually:
x ↦ (y ↦ x + y)
Now applying the function to two arguments means applying it twice:
Apply
(Apply
(FunDef "x"
(FunDef "y"
(BinOp (Var "x") PlusOp (Var "y"))))
(Num 10))
(Num 100)
This expression should eventually produce:
Num 110
First application¶
Substitute 10 for x:
FunDef "y"
(BinOp (Num 10) PlusOp (Var "y"))
Notice what the first application returns:
another function
That is exactly what currying says should happen.
Second application¶
Now apply that resulting function to 100:
BinOp (Num 10) PlusOp (Num 100)
and evaluate:
Num 110

This is not merely an analogy to Haskell currying.
We have now implemented the same structural idea inside our own language.
A Snag: Apply Does Not Always Contain a FunDef¶
Our first application rule was:
eval (Apply (FunDef varname body) arg) =
eval (substitute arg (Var varname) body)
Look carefully at the pattern:
Apply (FunDef varname body) arg
It only matches when the function position is already literally a FunDef.
But our curried two-argument expression has this outer shape:
Apply
(Apply ...)
(Num 100)
The first expression inside the outer Apply is itself another Apply.
It is not yet syntactically a FunDef.
So the old pattern does not match.
This is a great example of why thinking carefully about the structure of the abstract syntax matters. Our semantics should not demand that the function position already look like a function definition. It may be an expression that evaluates to one.

Fix: evaluate the function position first¶
Instead of pattern matching immediately on a FunDef, accept any function expression:
eval (Apply fexp arg) =
let FunDef varname body = eval fexp
in eval (substitute arg (Var varname) body)
The important new step is:
eval fexp
We first evaluate the function position.
We expect its result to be a function value:
FunDef varname body
Then we substitute the argument into that body.
This makes nested applications work.
Call-by-value version¶
If we choose call-by-value for the argument, the corresponding version is:
eval (Apply fexp arg) =
let FunDef varname body = eval fexp
val = eval arg
in eval (substitute val (Var varname) body)
Two separate decisions are happening:
- Function position: evaluate it until we obtain a function value.
- Argument position: choose when the argument should be evaluated according to our evaluation strategy.
That distinction will continue to matter as languages become more complicated.
The Interpreter So Far¶
Putting the pieces together, one version of our language now looks like this:
data Op
= PlusOp
| MinusOp
| TimesOp
| DivOp
deriving (Show, Eq)
type VarName = String
data Exp
= Var VarName
| BinOp Exp Op Exp
| Num Double
| FunDef VarName Exp
| Apply Exp Exp
deriving (Show, Eq)
Arithmetic operations:
evalOp :: Exp -> Op -> Exp -> Exp
evalOp (Num x) PlusOp (Num y) = Num (x + y)
evalOp (Num x) MinusOp (Num y) = Num (x - y)
evalOp (Num x) TimesOp (Num y) = Num (x * y)
evalOp (Num x) DivOp (Num y) = Num (x / y)
Evaluation:
eval :: Exp -> Exp
eval (Num x) =
Num x
eval (BinOp left op right) =
evalOp (eval left) op (eval right)
eval (FunDef varname body) =
FunDef varname body
eval (Apply fexp arg) =
let FunDef varname body = eval fexp
in eval (substitute arg (Var varname) body)
And substitution:
substitute :: Exp -> Exp -> Exp -> Exp
substitute _ (Var v) (Num x) =
Num x
substitute replacement (Var v) (Var w) =
if v == w
then replacement
else Var w
substitute replacement var (BinOp left op right) =
BinOp
(substitute replacement var left)
op
(substitute replacement var right)
substitute replacement var (FunDef varname body) =
FunDef varname
(substitute replacement var body)
substitute _ _ _ =
error "bad substitution"

This is already remarkable.
We have described:
- numeric values,
- arithmetic expressions,
- variable references,
- function values,
- function application,
- and an evaluation strategy,
using a small recursive datatype and a handful of recursive functions.
But We Are Only Halfway There¶
This module ends with an important warning:
We're actually only halfway there...
Our evaluator works for the examples we have chosen, but some deeper questions are now visible.
Consider:
FunDef "x"
(FunDef "x"
(BinOp (Var "x") PlusOp (Var "x")))
There are two different parameters named x.
What should happen if we substitute for the outer x?
Should substitution continue into the inner function body?
The simple substitution rule we wrote does:
substitute replacement var (FunDef varname body) =
FunDef varname (substitute replacement var body)
But that can ignore the fact that the inner FunDef "x" introduces a new binding for the name x.
Now consider:
FunDef "x"
(BinOp (Var "x") PlusOp (Var "y"))
Where does y come from?
It is not the function's parameter.
And finally:
How would we represent a recursive function?
Our FunDef stores a parameter and a body, but the function itself has no name. How could its body refer back to the function?

These are not small implementation annoyances.
They are clues that we need better concepts for understanding which variable occurrence refers to which binding.
That leads directly to:
- scope,
- free and bound variables,
- environments, and
- closures.
Those ideas let us handle names and functions much more systematically than naive textual substitution.
Pulling the Main Ideas Together¶
This module extends the "code is data" story from arithmetic to functions.
Function definitions are syntax¶
FunDef "x" body
is an ordinary Exp value representing a function.
Functions are values¶
Evaluating a function definition gives us the function itself:
eval (FunDef x body) = FunDef x body
This is the interpreter-level version of the principle from Module 02.2:
Functions are values.
Function application needs semantics¶
Apply f arg
does not explain itself.
We define its meaning through an evaluation rule.
Substitution gives us a first model of application¶
Applying:
x ↦ body
to:
arg
can be modeled as:
body[x → arg]
followed by evaluation.
Evaluation strategy is a language-design decision¶
Call-by-name and call-by-value differ only by a small change in our evaluator, but that small change affects:
- whether work is duplicated,
- whether unused arguments are evaluated,
- and even whether some programs terminate.
Currying falls naturally out of the representation¶
A multi-argument function becomes nested one-argument functions:
FunDef "x" (FunDef "y" body)
and multiple arguments become nested applications:
Apply (Apply f a) b
Our first substitution model is deliberately incomplete¶
Variable shadowing, free variables, and recursion expose weaknesses in naive substitution.
That is not a failure of the exercise. It tells us exactly what conceptual machinery we need next.
Where This Leaves Us¶
You should now be able to read expressions such as:
FunDef "x"
(BinOp (Var "x") PlusOp (Num 1))
and:
Apply
(FunDef "x"
(BinOp (Var "x") PlusOp (Num 1)))
(Num 5)
and explain both their syntax and their intended semantics.
You should also be able to explain:
e[x → e']
and trace how substitution recursively traverses an expression.
Most importantly, the ideas from the first several modules are now converging:
functions are values
+
recursive data structures
+
pattern matching
+
recursion
+
code is data
=
a small interpreter that represents and applies functions
The next step is to confront the problems we deliberately left unresolved: which binding does a variable refer to, what happens when names are shadowed, and what information must a function carry with it in order to behave correctly?
That takes us to scope, environments, and closures.
Finish the Module 04.2 Completion¶
The existing Module 4.2 materials identify four substantive completion questions, all of which appeared at the relevant points above:
- Describe what should happen when evaluating the provided application of the "plus one" function to
5. - Perform the four given substitutions.
- Complete the implementation of
substitute. - Compare the two proposed function-application evaluation strategies and discuss whether one can be better or worse.
Make sure you have submitted your responses to the Module 4.2 Completion assignment on Gradescope.