- From: Bethan Tovey-Walsh <bytheway@linguacelta.com>
- Date: Tue, 7 Jul 2026 11:50:21 +0100
- To: Steven Pemberton <steven.pemberton@cwi.nl>
- Cc: ixml <public-ixml@w3.org>
Sorry - I misspoke about lookbehind; you *can* do it, but not quite as I outlined it!
___________________________________________________
Dr. Bethan Tovey-Walsh
linguacelta.com
Golygydd | Editor geirfan.cymru
Croeso i chi ysgrifennu ataf yn y Gymraeg.
> On 7 Jul 2026, at 11:33, Bethan Tovey-Walsh <bytheway@linguacelta.com> wrote:
>
> Thanks! Very interesting.
>
> I think it would be a bit limiting to call this a "disambiguation" construct, to be honest. It provides lookahead and lookbehind*, which is a pretty major extension to what iXML can already do, as it makes the language context-sensitive.
>
> That being the case, I don't think I'd agree that it's the same kind of thing as the repetition operators. It doesn't really say "match zero of this"; it says "ensure that this does not match". We don't have any other such non-consuming match patterns in the language at present.
>
> *Although technically this implements only negative lookahead and lookbehind, you could also use it for the positive versions by doubling the operator.
>
> This:
>
> CAT: "cat", not_ER!.
> not_ER: "erpillar".
>
> means "match the string 'cat', as long as it's not followed by 'erpillar'" - i.e. find occurrences of "cat" that are not the beginning of the word "caterpillar".
>
> Then with this:
>
> CAT: "cat", not_ER!.
> not_ER: ER!
> ER: "erpillar".
>
> you can say "match 'cat', as long as it's followed by something that's not not "erpillar" - i.e. as long as it *is* the beginning of the word 'caterpillar'".
>
> You can also do lookbehind:
>
> CAT: not_COPY!, "cat".
> not_COPY!: "copy".
>
> So, "find occurrences of 'cat' that aren't part of the word 'copycat'".
>
> And you can do positive lookbehind as above, by doubling the negation.
>
> BTW
> ___________________________________________________
> Dr. Bethan Tovey-Walsh
>
> linguacelta.com
>
> Golygydd | Editor geirfan.cymru
>
> Croeso i chi ysgrifennu ataf yn y Gymraeg.
>
>> On 7 Jul 2026, at 11:08, Steven Pemberton <steven.pemberton@cwi.nl> wrote:
>>
>>
>>
>> On Tuesday 07 July 2026 00:38:11 (+02:00), Bethan Tovey-Walsh wrote:
>>
>> Thanks for sharing this, Steven.
>>
>> What are the changes to the grammar that describe this construct? And what exactly are its semantics? (From what you've described, I'm assuming that it's a kind of negative lookahead.)
>>
>> -term: factor;
>> option;
>> prohibition; {added}
>> repeat0;
>> repeat1.
>> prohibition: factor, -"!", s.
>>
>> The semantics are:
>>
>> In an alternative such as
>> A, B!, C
>>
>> if B succeeds (successfully parses) at that position, B! fails, and so the alternative fails; if B fails, B! succeeds, and C is started at the position B was started.
>>
>> Steven
>>
>>
>>
>> BTW
>>
>> ****************************************************
>> Dr. Bethan Tovey-Walsh
>> linguacelta.com
>> Golygydd | Editor geirfan.cymru
>> Croeso i chi ysgrifennu ataf yn y Gymraeg
>>
>>> On 6 Jul 2026, at 16:22, Steven Pemberton <steven.pemberton@cwi.nl> wrote:
>>>
>>> In the new implementation, we decided to use the notation "x!".
>>>
>>> The reasoning behind this is that all similar constructs in ixml are also postfix:
>>>
>>> x* zero or more
>>> x+ one or more
>>> x? zero or one
>>> x! zero
>>>
>>> and prefix characters are use for marks on serialisation.
>>>
>>> So here are some example disambiguation grammars.
>>>
>>> IDENTIFIERS AND KEYWORDS
>>> Each line is either an identifier or a keyword
>>>
>>> input: ((id; keyword), #a)*.
>>> id: keyword!, letter+, letter!.
>>> -letter: [L].
>>> keyword: ("if"; "then"; "else"), letter!.
>>>
>>> Input
>>> i
>>> if
>>> ifi
>>> the
>>> then
>>> thene
>>>
>>> Output
>>> <input>
>>> <id>i</id>
>>> <keyword>if</keyword>
>>> <id>ifi</id>
>>> <id>the</id>
>>> <keyword>then</keyword>
>>> <id>thene</id>
>>> </input>
>>>
>>> MIXED CHARACTERS
>>> Taken from the ixml tutorial, separating input characters into classes
>>> {example input:
>>> ...abcd()1234!!!}
>>> input: line+.
>>> line: (number; word; punc)*,
>>> -#a.
>>> number: digit+, digit!.
>>> word: letter+, letter!.
>>> punc:
>>> p+, p!.
>>> -letter: [L].
>>> -digit: ["0"-"9"].
>>> -p:
>>> [P].
>>> Output
>>> <input>
>>>
>>> <line>
>>> <punc>...</punc>
>>>
>>> <word>abcd</word>
>>>
>>> <punc>()</punc>
>>>
>>> <number>1234</number>
>>>
>>> <punc>!!!</punc>
>>>
>>> </line>
>>> </input>
>>> CODES
>>> Each line is either a header that ends with a code, or an item that is a series of words.
>>> catalogue: entry*.
>>> entry: header,
>>> item+.
>>> header: text, -" ", code, -#a.
>>>
>>> item: (text, " ")?, code!, word, -#a.
>>> -text: word**"
>>> ".
>>> -word: (l; d)+, (l;d)!.
>>> @code: l, l,
>>> l, d, d, d.
>>> -l: [L].
>>>
>>> -d: ["0"-"9"].
>>> Input
>>> Fiction
>>> fic001
>>> Brave New World
>>> 1984
>>> NonFiction non123
>>> Translating
>>> Beaudelaire
>>> The Sixth
>>> Extinction
>>>
>>> Output
>>> <catalogue>
>>>
>>> <entry>
>>> <header
>>> code="fic001">Fiction</header>
>>> <item>Brave
>>> New World</item>
>>>
>>> <item>1984</item>
>>> </entry>
>>>
>>> <entry>
>>> <header
>>> code="non123">NonFiction</header>
>>>
>>> <item>Translating Beaudelaire</item>
>>>
>>> <item>The Sixth Extinction</item>
>>>
>>> </entry>
>>> </catalogue>
>>> Steven
>
>
>
Received on Tuesday, 7 July 2026 10:50:40 UTC