Pep & Nom

home | Syntagma | docs | examples | translators | download | journal | blog | all blog posts

If you want to learn how to do something, then watch the animals and insects, and learn from them. Kogi Saying

(definitely not) frequently asked questions about pep & nom

A set of questions and answers about the PEP virtual machine and the NOM scripting language.

how do you know that nom can parse context-sensitive languages?

Because it can recognise and parse the language {a^n b^n c^n}

can pep/nom do indent parsing like in python?

Yes, but I haven't actually written an example script which demonstrates this.

can pep/nom do type checking and type inference?

Yes, I think so, but the type checking or inference needs to be in a separate “pass” of the compiler or translator. I haven't yet written an example script for this, but it is on my list of priorities. The technique for type checking is to store the type of the variable or expression in a marked tape cell, where the mark is the name of the variable (with, perhaps the scope of the variable as a prefix). One challenge with this approach is to allocate sufficient tape cells (at the top of the tape) to accommodate all the variables that will be encountered. It occurred to me to implement a “reserve” nom command that will reserve a number of cells at the top of the tape and set a base variable which will ensure that a pop command will not decrement the tape pointer when the tape pointer is at the base.

why is there no "contains" test in the nom language?

The nom language contains begins-with (eg B"abc") and ends-with (eg E"abc") tests that check if the workspace text buffer ends with or starts with the given text (see tests ). So why is there no contains test?

Well, to answer in a round-about way, there are many functions to manipulate the workspace buffer that may be useful. And a contains test would be handy in some circumstances, but my aim in designing the Pep & Nom system was to keep the virtual machine and language as simple and essential as possible.

The Pep & Nom system is based on the Unix filter model which means that a nom script is designed to “edit” a text stream (or parse and translate it). This means that it should (almost) always be possible to catch language patterns with the begins-with and ends-with tests.

can i use nom to write a "pretty printer" for a language ?

Yes, it’s simple because pretty-printers (or source code colourisers) don’t (normally) actually parse the source code. They only tokenise which is to say, they recognise pattern like quoted text, command names, punctuation etc and format them. An example of a pretty printer in Nom is /eg/nom.tohtml.pss or /eg/nom.tolatex.pss which does the same thing but into LATEX not HTML .

If you want to write something that properly indents and formats the source code then you will need to actually parse the language ( NOM can do that too). Possibly the most difficult task is then deciding where to split lines that are too long. But you are not flying an æroplane, just do something that doesn't look too bad.

how do i install pep/nom ?

Follow the instructions at www.nomlang.org/download.

can nom be used to create a modern programming language?

I believe so. It can certainly be used to parse and compile context-free languages but I have not yet written a successful type checker using nom. But it is definitely possible, especially using the nom commands mark and go in their non-parameter versions (this essentially allows the PEP tape structure to be used as an en.wikipedia.org/wiki/associative_array )

is pep/nom in active development?

I have been developing, or trying to develop Pep & Nom for a long time. I 1st thought of the idea several decades ago, but judging by my notes and work-diaries I didn't really achieve a decent interpreting implementation until around 2019. Since then I have worked off and on on the system. Sometimes I take long breaks from it to do other things but when I come back to it, it always seems very interesting and powerful.

So the answer is “yes” (as of 2026).

is the pep/nom system widely used?

No, it isn't. I believe that almost nobody knows about it, which is why I am writing this website and also this FAQ . I believe the system is extremely interesting and potentially very useful and what is more, its free and open source

what can i use the nom language for?

You can use it for parsing, translating, compiling or transpiling context-free (and some context-sensitive) languages such as Lisp, JSON data, CSS code, palindromes, regular expressions, CSV data, XML and so on and so forth. You can use it to create your own mini-languages. You can use it to explore and understand the grammar of languages and patterns without having to code in a full computer language. Nom and Pep can also handle certain types of context-sensitive languages or patterns ( [XML] maybe context-sensitive)

who is noam chomsky?

A linguist who came up with an interesting classification of formal languages.

why is the nom language so cryptic?

The Nom language closely reflects the virtual machine which it manipulates. This machine has a stack a tape (an array with a pointer) and a workspace buffer as well as some other minor registers. So, for example, the ++ command increments the pointer to the tape element and the push command pushes one parse token from the (beginning of) the workspace onto the (top of the ) stack. So once you are familiar with the virtual machine you will find the language not so cryptic.

Having said that, it was actually my intention, when I first thought of pep&nom that I would write a more expressive and natural language “on top” of it. In other words, I would use nom to write a compiler for a language which would reflect more closely a kind of BNF grammar.

The page /doc/syntax/nom.syntax.html contains an example of compiling a simple (toy) BNF format into NOM which can be used as the basic of a more expressive language that compiles into the nom script language.

In my defence, I could say that all languages or mini-languages that are based on virtual machines tend to be cryptic, for example SED AWK FORTH etc.

what tools did you use to develop pep/nom?

I work mostly on Linux or MacOS so the tools I use are available on those machines such as SED VIM the c language and the GCC compiler (for the pep interpreter ) although the PEP interpreter also compiles with TCC and should compile with any standard c compiler.

there are already too many computer languages. why another one?

Nom is not a general purpose computer language, it is a “domain-specific ” computer language designed to recognise and 'translate' text patterns (more formally called context-free languages).

Nom is also important because it highlights that a compiler it really just a map function from one 'language' to another. A formal language is a set of strings in a given alphabet. So a compiler (if we ignore the final stage of converting to binary code) is a map function from one set of strings to another.

a compiler as a map

  language A: { 'and','with','for','to','from' }
    --> (maps to)
  language B: { 'y','con','para','a','de' }
 

But, of course the language map function does not have to be one-to-one. When compilers are written as computer code, this map function is hidden or obfuscated . But when the compiler is written as a text-based map function, we are able to “multiply” map functions (that is “chain” compilers together to create new compilers). Nom already does this chaining with it’s translation scripts

The nom language is still code-oriented rather than data-oriented but it is an important step towards regarding compilers as language map functions.

how many people use nom?

As far as I know, in September 2026, only me.

is nom a commercial project?

No, it is an idea, that I feel could be useful and interesting in many software circumstances. I believe, for example, that it could be used to write very small simple compilers for microcontrollers (or other 'small' machines) that would be capable of running on the microcontroller like FORTH or micro-python.

what is the "workspace* ?

This is the central text buffer of the PEP virtual machine Each part of the machine is documented in /doc/machine/ and the workspace buffer is documented at workspace and the parse-token stack is documented at stack

why is the "workspace" called the "workspace"?

Because, like a number of things in the Pep & Nom system, it was influenced by the SED stream editor and I believe that the GNU sed program uses similar terminology.

how is pep/nom different from sed ?

Sed reads the text input-stream line by line and uses regular expression patterns to transform (or edit ) the text stream. NOM can read the input stream character-by-character or any other way you wish and uses grammar rules to transform (compile/transpile/translate) the input-stream.

can linguists use nom?

I think so, but it is not really designed (at the moment) for translating human languages. But it may be useful for linguists to understand how formal grammars function (and their limitations).

There are many complexities to parsing human languages (for example the nexus between semantic considerations and grammatical rules) but one of the more obvious ones is the “ambiguity” of certain sentences: that is, one sentence may have several different possible ways to parse, or initial parsings. Since Pep & Nom cannot “backtrack” in the input stream it cannot choose between multiple parsings.

are there any bugs in pep/nom?

Yes, unfortunately there will be. This is also because I have been working on this by myself off and on. I just found a bug in the mark and go code where more than one tape cell is marked with the same 'tag' . But generally the system works remarkably well. For a list of known bugs, see the document /doc/pepnom.doc.bugs.html

why is writing nom like playing blindfold chess?

The reason is that you need to maintain the state of the PEP machine in your head as you are manipulating it. That is similar to the way that I blindfold chess player maintains the position of the board in his or her head during the game. But with nom you have several tools to 'peek' into the state of the machine during the development and execution of a script. These tools are outlined in the document /doc/howto/nom.debug.html . One simple one is the state command which prints the state of the machine at the time that the command is executed.

can i use regular expressions in nom?

No Context-free and context sensitive languages are supersets of regular languages, so Nom can recognise and translate regular languages but there is no regular-expression engine built into nom. There are a couple of main reasons for this. Firstly, it reduces the complexity of the code required to implement Nom, and the pep virtual machine and allows Pep & Nom to run on small microcontrollers. Secondly, the aim of Nom is to allow recognising and translating language patterns without limiting the analysis to regular languages.

have the large language models made pep/nom obsolete?

I don't think so, although since Nom is not used anyway the question is moot. A nom script can be thought about as a compiled representation of an LLM response and nom's approach to pattern recognition may actually be complementary to the LLM approach. However, a better response to this question probably needs to await the further development and use of the LLMs.

can i compile a nom script into an executable?

Yes, this is done via one of the translation scripts in the www.nomlang.org/tr/ translation folder. The better, and more recently written translation scripts are called “nom.to<lang>.pss ” such as /tr/nom.torust.pss or /tr/nom.todart.pss or /tr/nom.toperl.pss A translation script is a nom script that translates nom scripts into other (computer) languages. So the script /tr/nom.tolua.pss translates a nom script into the lua language.

Currently the best (least buggy) option for translating a nom script into an executable would be via rust or dart

what is syntagma?

Syntagma is a new (2026) implementation of a parsing language which is compiled (by nom) into a nom script. The syntax of syntagma is much closer to the extended backus-naur form (EBNF) and should be more intuitive to use for people who have worked with formal grammars. Because syntagma is a “higher level” language, syntagma scripts tend to be much more succinct than the equivalent nom script.

nom scripts appear ridiculously verbose, considering the actual

task they are carrying out. Why?

This is true, especially when comparing the nom script to an equivalent regular expression. This could be mitigated by using the one-character command equivalents for nom, but the scripts would still be more verbose. Compared to regular-expressions, nom is able to recognise and compile a wider range of patterns and is able to translate itself to other languages, which may justify the extra verbosity.

why does the pep interpreter throw segmentation faults?

Because I haven't debugged it properly. Sometimes it works and sometimes it doesn't and recently I have been putting more time and effort into the implementation of syntagma and the translation scripts. Given a good translation script, which has translated itself, the pep interpreter can actually be ignored completely. Examples of self-translation can be seen in the /tr/ folder. The rust source file /tr/nom.torust.rs is the source of a rust program which has been generated by running the nom rust translator on itself as follows

 pep -f nom.torust.pss nom.torust.pss > nom.torust.rs

This can then be compiled to a executable ( /tr/nom.torust.exe ) which can act as a replacement for the pep interpreter. This process can then self replicate as follows (even though I can't think of any practical purpose for this).

 cat nom.torust.pss | ./nom.torust.exe > nom.torust.2gen.rs

An interesting corollary is that diff nom.torust.2gen.rs nom.torust.rs returns no output, meaning that the 2 rust source files are the same.