@Doc

Inline Syntax Specification

@Doc Inline Syntax Specification v1.4

0. Table of Contents


1. Design Philosophy

@Doc adopts:

Only Known Commands Trigger Parsing

Only known commands carry grammatical meaning.

Unknown commands are always treated as plain text.

The goals of this design:

  • Lower the learning cost
  • Avoid conflicts with email and mention systems
  • Improve AI parsing stability
  • Improve editor fault tolerance
  • Preserve the DSL's extensibility
  • Establish a stable and predictable AST

2. Lexer Behavior Definition

Command Parsing Rules

When the Lexer scans an @, it should process it in the following priority order:

  1. If followed by @@
    • Parse it as a single literal @ character
  2. If what follows matches a registered command name
    • Enter the corresponding grammar-parsing flow
  3. If it matches no known command
    • Output the entire span as plain text

Examples

InputResult
@mark[hello]Parsed as a mark node
@@markOutputs @mark
[email protected]Plain text
@GitHubPlain text
@unknownPlain text

3. Ambiguity Resolution Rule

Since @Doc adopts:

Known Command Recognition

the Lexer must first attempt to recognize known commands before falling back to plain-text mode.

In other words:

inline-node takes precedence over plain-text-char.

The Lexer must follow:

Starts with @
↓
Is it @@ ?
↓
Does it exist in the Command Registry ?
↓
Yes → Inline Node
No → Plain Text

Therefore:

@mark[hello]

must be parsed as:

InlineNode(mark)

rather than:

Text('@')
Text('m')
Text('a')
Text('r')
Text('k')
...

4. Complete EBNF Grammar Definition

(* ==========================================================================
   Entry Point
   ========================================================================== *)

inline-stream =
    { inline-node | plain-text-char } ;

inline-node =
      mark
    | color
    | bordered
    | bold
    | italic
    | underline
    | del
    | raw
    | sup
    | sub
    | fn
    | defn
    | kbd
    | link
    | br
    | escape ;

(* ==========================================================================
   Inline Nodes
   ========================================================================== *)

mark      = "@mark" , [ styles ] , content ;
color     = "@color" , [ styles ] , content ;
(* @bordered shares @color's exact {styles} slot and swatch (see §7),
   applied as a text border instead of a foreground color. *)
bordered  = "@bordered" , [ styles ] , content ;
bold      = ( "@bold" | "@b" ) , content ;
italic    = ( "@italic" | "@i" ) , content ;
underline = ( "@underline" | "@u" ) , content ;
del       = "@del" , content ;

raw       = "@raw" , raw-content ;

sup       = "@sup" , content ;
sub       = "@sub" , content ;

(* Footnotes:
   fn   = the in-text reference marker (superscript), carries only the number
   defn = the footnote definition body, carries the number and the actual content
*)
fn        = "@fn" , "[" , integer , "]" ;
defn      = "@defn" , modifier , content ;

kbd       = "@kbd" , "[" , key , "]" ;

link      = "@link" , uri , content ;

br        = "@n" ;

escape    = "@@" ;

(* ==========================================================================
   Shared Components
   ========================================================================== *)

content =
    "[" ,
        { content-element } ,
    "]" ;

content-element =
      inline-node
    | plain-text-char ;

(* The actual termination rule for raw-content is "bracket-depth counting,"
   not "reaching the first unescaped ]" — balanced-bracket-group uses a
   recursive production to express "as long as brackets are paired inside,
   they can nest freely with no escaping needed at all"; only truly
   unpaired brackets need to be escaped. See 9. @raw Opaque Domain for
   details. *)
raw-content =
    "[" , { raw-unit } , "]" ;

raw-unit =
      escaped-at-close-bracket   (* "@@]" → literal "@]" *)
    | escaped-at-open-bracket    (* "@@[" → literal "@[" *)
    | escaped-close-bracket      (* "@]"  → literal "]" (only used for an unpaired ]) *)
    | escaped-open-bracket       (* "@["  → literal "[" (only used for an unpaired [) *)
    | balanced-bracket-group     (* paired, nestable literal brackets, content unrestricted *)
    | raw-char ;

balanced-bracket-group =
    "[" , { raw-unit } , "]" ;

escaped-at-close-bracket = "@@]" ;
escaped-at-open-bracket  = "@@[" ;
escaped-close-bracket    = "@]" ;
escaped-open-bracket     = "@[" ;

raw-char =
    any-unicode-char - "]" - "[" ;

uri =
    "(" ,
        { text-char - ")" } ,
    ")" ;

modifier =
    "(" ,
        { text-char - ")" } ,
    ")" ;

(* Additional Lexer restriction: `text-char` itself includes newlines, so a
   literal reading would mean "an unclosed "{" can swallow everything all
   the way to any "}" later in the document" — a "{" the author is still
   typing would swallow every node in between (e.g. everything before the
   curly brace in the @code block below) whole into styles, silently
   vanishing from the AST. The actual semantics of styles are a short,
   comma-separated list of tokens with no example ever spanning multiple
   lines, and the editor's Monarch rule (/\{[^}]*\}/, matched line by line)
   doesn't support spanning lines either — so the Lexer stops scanning at
   whichever of "}", end of line, or "[" (start of a content slot) comes
   first; both end-of-line and "[" are treated as unclosed. See
   scanStylesEnd() in src/Lexer.ts. *)
styles =
    "{" ,
        { text-char - "}" - newline - "[" } ,
    "}" ;

key =
    { text-char - "]" } ;

(* @color's semantic constraint on its {styles} content — see §7 for the
   full validation rule (must match /^#[0-9a-fA-F]{6}$/); the terminal itself
   is grammar-level only, exact digit-count/case validation is semantic-level,
   same split as `styles` below. *)
hex-color =
    "#" , hex-digit , hex-digit , hex-digit , hex-digit , hex-digit , hex-digit ;

hex-digit =
      digit
    | "a" | "b" | "c" | "d" | "e" | "f"
    | "A" | "B" | "C" | "D" | "E" | "F" ;

integer =
    digit ,
    { digit } ;

digit =
      "0" | "1" | "2" | "3" | "4"
    | "5" | "6" | "7" | "8" | "9" ;

(* ==========================================================================
   Character Sets
   ========================================================================== *)

(* Note:
   plain-text-char has lower precedence than inline-node.

   The lexer MUST always attempt known command recognition
   before falling back to plain text.
*)

plain-text-char =
    any-unicode-char ;

text-char =
    any-unicode-char ;

letter =
    Unicode Letter ;

symbol =
    Unicode Symbol ;

5. Escape Rule

Syntax

@@

Output

@

Purpose

Used when the user needs to output a syntax keyword itself.

This is a global escape rule, applicable in the general inline-stream context. @raw has its own independent escaping rules internally; see 9. @raw Opaque Domain.


Examples

Input:

@@mark

Output:

@mark

Input:

@@bold[hello]

Output:

@bold[hello]

Input:

Email: test@@example.com

Output:

Email: [email protected]

Although this form is valid, since:

example

is not a known command, in practice you can simply write:

Email: [email protected]

without needing to escape it.


6. Unknown Command Fallback

If what follows @ is not a valid command name, the parser must fall back to plain-text mode.

Example:

@github

Output:

@github

[email protected]

Output:

[email protected]

@my_custom_tag

Output:

@my_custom_tag

This rule effectively prevents conflicts with:

  • Email
  • Social media handles
  • Discord Mention
  • GitHub Username
  • Chat Mention System

from occurring.


7. @mark / @color / @bordered Styles Semantics

@mark supports an optional styles modifier syntax:

@mark{style}[content]

where:

  • style is a comma-separated style token string (style token list).
  • content is the text content being marked.
  • styles is optional syntax; when omitted, it is equivalent to a plain highlight mark:
@mark[important content]

Style Token Semantics

The content of style is a comma-separated Color Token string, representing the highlight color; the renderer maps it semantically to an actual color value. Two forms are supported, either of which may be used:

  • Named tokens (the renderer defines the actual color values):
    yellow / red / green / blue / orange / purple / gray
  • Hexadecimal tokens (starting with #, followed by 6 hex digits, case-insensitive; the renderer MUST use the specified value directly and MUST NOT remap it):
    #ff0000 / #3366FF / #00c896
    A token that does not match /^#[0-9a-fA-F]{6}$/ (e.g. #f00, #gggggg) is not considered a valid hex token, and follows the general fault-tolerant spirit of Unknown Command Fallback (see Renderer Behavior below).

Changelog: An earlier version separately defined three modifier tokens, underline/strikethrough/bordered, which have been removed. underline duplicated the semantics of the @underline node; strikethrough duplicated the semantics of the @del node; bordered has been promoted to its own independent node, @bordered (see below). style now only carries color semantics, and no longer mixes in modifier semantics.


Examples

@mark[default highlight]
@mark{yellow}[yellow highlight]
@mark{red}[red highlight]
@mark{#3366ff}[hex background color]

@color — Changing Text Color

@mark changes the background (highlight) and cannot change the color of the text itself. @color fills this gap:

@color{#ff0000}[This text is red]

@color shares the same {styles} field as @mark (see the EBNF above), and is itself optional — when omitted, the renderer falls back to a default color, behaving the same way @mark[content] does when {styles} is omitted.

@color{blue}[This text is dark blue]

@color accepts the same seven named color tokens as @mark (yellow/red/green/blue/ orange/purple/gray), and also accepts a single hexadecimal token (/^#[0-9a-fA-F]{6}$/). Both share the same set of token names syntactically, but **the actual color values they map to are independent**: @mark's color scale is tuned for light highlight backgrounds, and using it directly as a text foreground color would have insufficient contrast and be hard to read, so renderers typically maintain a separate, darker-toned lookup table specifically for @color (rather than reusing @mark's). The renderer MUST ignore malformed or unrecognized values and fall back to some default value, rather than throwing an error:

@color{not-a-color}[This text has no color specified, and gracefully falls back to the default color]

@bordered — Text Border

@bordered adds a border around text, sharing exactly the same {styles} field as @color — the same braces, equally optional, the same seven named tokens plus hex, and the same color-swatch lookup table (in implementation, @color's resolver can be reused directly) — the only difference is that it's applied to the border rather than the text color:

@bordered[default border]
@bordered{blue}[blue border]
@bordered{#3366ff}[hex border color]

When {styles} is omitted or given an unrecognized value, the renderer falls back to its default border style rather than throwing an error, echoing the fault-tolerant spirit of §6 Unknown Command Fallback. This node replaces the bordered modifier token that previously lived in @mark's {styles}, becoming its own independent first-class node, playing the same role as underline (now @underline) and strikethrough (now @del).


Renderer Behavior

  • The renderer MUST at least support the default highlight style (@mark) / default border style (@bordered) when styles is omitted.
  • The renderer MAY decide for itself the actual color value each named color token maps to (e.g. yellow may differ between dark mode and light mode; @mark and @color/@bordered may, and usually should, each maintain their own separate lookup tables, for the reasons given above); a hexadecimal token, however, MUST be used directly as specified and MUST NOT be remapped.
  • The renderer MUST ignore unrecognized tokens (including malformed hex tokens), and SHOULD fall back to some default value rather than throwing an error — this behavior is consistent with the fault-tolerant spirit of "unknown commands fall back to plain text" in 6. Unknown Command Fallback, but its scope is limited to inside styles; @mark[content]/@color[content]/@bordered[content] themselves are still parsed normally as their corresponding nodes. The exact shape of the fallback is up to the renderer to decide: it could be "no extra color" (falling back to the default text color/border color), or it could reuse @mark's default highlight color when styles is omitted (since all three share the same field shape, this is visually natural).
  • The delimiter between tokens is a fixed half-width comma ,, with any amount of whitespace allowed before and after (the Parser should auto-trim). This rule applies to @mark's styles; the {} of @color/@bordered only allows a single token (a color token or a hex value) and does not use comma-separation.

EBNF Supplementary Notes

Corresponding to the following in 4. Complete EBNF Grammar Definition:

(* Additional Lexer restriction: `text-char` itself includes newlines, so a
   literal reading would mean "an unclosed "{" can swallow everything all
   the way to any "}" later in the document" — a "{" the author is still
   typing would swallow every node in between (e.g. everything before the
   curly brace in the @code block below) whole into styles, silently
   vanishing from the AST. The actual semantics of styles are a short,
   comma-separated list of tokens with no example ever spanning multiple
   lines, and the editor's Monarch rule (/\{[^}]*\}/, matched line by line)
   doesn't support spanning lines either — so the Lexer stops scanning at
   whichever of "}", end of line, or "[" (start of a content slot) comes
   first; both end-of-line and "[" are treated as unclosed. See
   scanStylesEnd() in src/Lexer.ts. *)
styles =
    "{" ,
        { text-char - "}" - newline - "[" } ,
    "}" ;

At the lexical level, styles itself is only defined as "an arbitrary character sequence wrapped in curly braces"; the actual token splitting (comma-separated, recognizing color tokens and modifier tokens) belongs to semantic-level processing — it is not the grammar-level responsibility of the Lexer/Parser, but is left to the renderer or a later semantic analysis stage. This design ensures that:

  • Adding new style tokens (e.g. italic, bold in the future) does not require modifying the EBNF grammar definition itself.
  • Different renderers can extend or trim their supported token set on their own, consistent with the "preserve the DSL's extensibility" goal in 1. Design Philosophy.

@link accepts any valid URI or URI-like identifier.

@link(uri)[content]

where:

  • uri is the identifier of the target resource.
  • content is the display text.

Renderer URI Inference

The renderer MAY automatically infer the URI scheme based on the content of uri.

For example:

InputActual renderer URI
@link(example.com)[Official Website]https://example.com
@link([email protected])[Contact me]mailto:[email protected]
@link(+886912345678)[Customer Service Phone]tel:+886912345678

If uri already explicitly specifies a scheme:

@link(https://example.com)[Official Website]
@link(mailto:[email protected])[Contact me]
@link(tel:+886912345678)[Customer Service Phone]

The renderer MUST use the specified value directly, and MUST NOT infer or modify it.


Supported URI Examples

The following are all valid uri values:

https://example.com
mailto:[email protected]
tel:+886912345678
ftp://example.com/file.zip
discord://channel/123
vscode://file/path
file:///tmp/test.txt

@Doc itself does not restrict the type of URI.

The actual support for a given URI is determined by the renderer.


9. @raw Opaque Domain

@raw belongs to:

Opaque Domain

Once the parser enters:

@raw[

afterward:

  • No internal syntax is parsed.
  • All keywords such as @mark, @bold, @link, etc. are treated as plain text.
  • The global @@ escape rule (5. Escape Rule) no longer applies; the raw domain has its own independent, local set of rules (see below).

Termination Rule: Bracket-Depth Counting, Not "Reaching the First ]"

The implementation model for @raw[...] is bracket-depth counting, and this must be made clear up front, because it directly determines when the escaping rules should and should not be used:

  • [ increases the depth by +1, ] decreases it by -1; the ] that brings the depth back to zero is the true end.
  • In other words, as long as the brackets are paired, you can copy them verbatim, with no escaping needed at all — the @mark[hello] inside @raw[@mark[hello]] itself has paired left and right brackets, so the Parser correctly ends at the outermost ], outputting @mark[hello] as literal text.
  • The escaping rules exist specifically for unpaired brackets — for example, when you just want to write a single, standalone literal ], or quote a content fragment whose own brackets are unbalanced. If you add unnecessary escaping to a bracket pair that is already balanced (e.g. writing the end of @mark[hello] as @mark[hello@]), the ] consumed by the escape does not bring the depth count back to zero, so the +1 depth caused by the earlier [ in @mark[ can never find a matching ] to cancel it out — the Parser can only keep searching forward, swallowing more and more of the outer content (potentially the entire document) before it finally errors out. Write balanced brackets as-is; always escape unbalanced brackets; escape characters do not participate in depth counting.

The escaping rules are symmetric — both ] and [ each have a "single-character escape" and an "escape the @ itself followed by that character" form, four rules in total; when scanning, they are matched in the following priority order (longer, more specific sequences first):

PriorityInputOutputDescription
1@@]@]Outputs the two literal characters @]
2@@[@[Outputs the two literal characters @[
3@]]Outputs an unpaired literal ] (does not affect depth counting)
4@[[Outputs an unpaired literal [ (does not affect depth counting)

Examples

Input:

@raw[@mark[hello]]

Output:

@mark[hello]

Explanation: @mark[hello] itself has paired brackets, so the depth count goes 1→2→1, and only the outermost ] brings the depth back to zero and ends @raw. No escaping is needed at all — this is the most common usage (demonstrating a complete, bracket-balanced piece of @Doc syntax inside raw content).


Input:

@raw[@@]

Output:

@@

Explanation: here the ] immediately following @@ is the closing bracket of raw-content itself, so the @@] special case is not triggered (@@] must be three consecutive characters); it should be parsed as: the literal characters @@ (output as-is, since the global escape rule is disabled) + the closing ].


Input:

@raw[Today I'm afraid the @] will be detected]

Output:

Today I'm afraid the ] will be detected

Explanation: here @] is an unpaired literal ] (there is no matching [ before it), so it must be escaped — otherwise it would be treated as the end of @raw itself, causing the following "will be detected" to fall outside the raw content.


Input:

@raw[Today I'm afraid the @@] will be detected]

Output:

Today I'm afraid the @] will be detected

Input (escaping used where it shouldn't be — counter-example):

@raw[Here, @mark[hello@] stays as-is]

Rendering semantics: inline code

The rules above define Parser behavior (content is not parsed; it is kept verbatim). The corresponding rendering semantics are inline code — the same thing Markdown's backticks `code` express:

  • A Renderer SHOULD present @raw as a monospace inline element. The HTML Route uses <code>.
  • This is distinct from @code, which is a block (<pre><code> in HTML); @raw is inline.
  • It is also distinct from @kbd, which denotes a physical key and conventionally carries a border and key-cap look; @raw is code text and needs only a monospace font and a tint.
  • This does not license a Renderer to parse the content — adding rendering semantics does not change the opaque domain's parsing rules.

Why @raw rather than a separate node: raw-escaped is the only content mode in the language with a real escape mechanism (@] / @[ plus bracket-depth counting), and "the content must not be parsed" is precisely the defining requirement of inline code. The two use cases overlap almost entirely — every example above in this section is showing code.


10. Nested Parsing

Because:

content-element =
      inline-node
    | plain-text-char ;

@Doc supports full recursive nesting.

For example:

@bold[
    This is bold text,
    inside there is
    @mark{yellow}[an important highlight]
    and
    @underline[an underline]
]

Its AST structure is:

Bold
├── Text
├── Mark
└── Underline

11. Parser Recovery Strategy

When the Parser encounters an unclosed structure:

@bold[hello

or:

@mark{red}[hello

Two suggested modes are provided:

Strict Mode

Throws a syntax error directly:

Unexpected EOF while parsing @bold

Editor Mode

Allows the editor to automatically complete missing closing symbols:

]

to improve the real-time editing experience.


12. Architecture

Recommended parsing pipeline:

Source Text
    ↓
Lexer
    ↓
Token Stream
    ↓
Parser
    ↓
AST
    ↓
Renderer

The renderer can freely output:

  • HTML
  • React
  • PDF
  • DOCX
  • Markdown
  • Discord
  • Terminal
  • Custom UI

13. Core Principle

The core goal of @Doc is not to replace Markdown.

but rather to establish:

Human Editable Machine Deterministic AI Friendly Cross Platform

a next-generation document interchange format.


14. Simplified Syntax Aliases

@bold/@italic/@underline provide the simplified aliases @b/@i/@u — purely a shorthand at input time. The Parser normalizes it to the canonical name before creating the AST node (node.type is always the canonical name); the renderer never needs to, and never does, distinguish which form the author actually typed.

CanonicalAlias
@bold@b
@italic@i
@underline@u

(The Block Syntax aliases @h/@p for @heading/@paragraph are defined in Block Syntax Specification §11.)