Re: [whatwg/url] Add URL path equality and prefix comparison (PR #929)

@domenic commented on this pull request.

This all feels a little unclean from the pure URL Standard perspective, but I appreciate this is needed for cookies and you're trying to balance.

No really substantial comments, although I think my final comment on the note is worth acting on.

> @@ -3220,15 +3220,18 @@ these steps. They return an <a>ASCII string</a>.
 
 <div algorithm>
 <p>The <dfn export lt="URL path serializer|URL path serializing">URL path serializer</dfn> takes a

This feels subjectively icky. I'd rather have two algorithms which take well-defined types, with one delegating to the other. It'd be a bonus if the one that takes a path could stay non-exported.

> @@ -3257,6 +3260,54 @@ run these steps:
 </ol>
 </div>
 
+<div algorithm="path equal">
+<p>To determine whether a <a for=/>URL path</a> <var>A</var>
+<dfn export for="URL path" lt=equal>equals</dfn> <a for=/>URL path</a> <var>B</var>, run these steps.
+They return a boolean.
+
+<ol>
+ <li><p>If <var>A</var> is a <a for=/>URL path segment</a> or <var>B</var> is a
+ <a for=/>URL path segment</a>, then return whether <var>A</var> is <var>B</var>.

Might be nice to link to `[=string/is=]`.

> @@ -3257,6 +3260,54 @@ run these steps:
 </ol>
 </div>
 
+<div algorithm="path equal">
+<p>To determine whether a <a for=/>URL path</a> <var>A</var>
+<dfn export for="URL path" lt=equal>equals</dfn> <a for=/>URL path</a> <var>B</var>, run these steps.
+They return a boolean.
+
+<ol>
+ <li><p>If <var>A</var> is a <a for=/>URL path segment</a> or <var>B</var> is a
+ <a for=/>URL path segment</a>, then return whether <var>A</var> is <var>B</var>.
+
+ <li><p>If <var>A</var>'s <a for=list>size</a> is not <var>B</var>'s <a for=list>size</a>, then
+ return false.
+
+ <li><p><a for=list>For each</a> <var>index</var> of <var>A</var>'s <a for=list>indices</a>: if

I guess https://github.com/whatwg/infra/pull/710 would be useful here, but stalled.

> +
+ <li><p><a for=list>For each</a> <var>index</var> of <var>A</var>'s <a for=list>indices</a>: if
+ <var>A</var>[<var>index</var>] is not <var>B</var>[<var>index</var>], then return false.
+
+ <li><p>Return true.
+</ol>
+
+<p class=note>Comparing <a for=/>URL path segments</a> rather than serializations matters for
+<a for=/>URL paths</a> that were not produced by the <a>URL parser</a>: the serialization of an
+<a for=url>opaque path</a> can be identical to that of a <a for=/>list</a> of
+<a for=/>URL path segments</a>.
+</div>
+
+<div algorithm="path starts with">
+<p>To determine whether a <a for=/>URL path</a> <var>A</var>
+<dfn export for="URL path" lt="start with|starts with">starts with</dfn> <a for=/>URL path</a> <var>B</var>, run

>100 chars

> + <li><p>If <var>A</var> is a <a for=/>URL path segment</a> or <var>B</var> is a
+ <a for=/>URL path segment</a>, then return false.
+
+ <li><p>If <var>B</var>'s <a for=list>size</a> is greater than <var>A</var>'s
+ <a for=list>size</a>, then return false.
+
+ <li><p><a for=list>For each</a> <var>index</var> of <var>B</var>'s <a for=list>indices</a>: if
+ <var>A</var>[<var>index</var>] is not <var>B</var>[<var>index</var>], then return false.
+
+ <li><p>Return true.
+</ol>
+
+<p class=note>As <a for=/>URL path segments</a> never contain U+002F (/), this can only be true when
+<var>B</var> covers whole segments of <var>A</var>. That makes it suitable for deciding whether one
+<a for=/>URL path</a> is contained by another, which comparing serializations is not: the
+serialization of « "<code>foo</code>" » is a prefix of that of « "<code>foobar</code>" ».

This note is helpful, but still a little confusing, in part because of how it moves between "starts with", "is a prefix", and "contains". Suggested, including some notable cross-links:

> As URL path segments never contain U+002F (/), A starts with B only when B covers whole segments of A. This makes it suitable for deciding whether one URL path is "contained" by another. Comparing [=URL path serializer|serializations=] is not appropriate: the serialization of « "foobar" » [=string/starts with=] the serialization of « "foo" », but conceptually, a URL with the former path does not contain a URL with the latter path.

This may further benefit from a second note, or maybe warning, explaining that the concept of paths "containing" each other is not really a good precedent, and is only used by legacy APIs like cookies (or service worker?) to impose a kind of path-based authority system that is discordant with the same origin policy used by the rest of the platform.

-- 
Reply to this email directly or view it on GitHub:
https://github.com/whatwg/url/pull/929#pullrequestreview-5096675007
You are receiving this because you are subscribed to this thread.

Message ID: <whatwg/url/pull/929/review/5096675007@github.com>

Received on Thursday, 3 September 2026 01:03:35 UTC