linuxmemes

20756 readers

187 users here now

I use Arch btw

Sister communities:

LemmyMemes: Memes
LemmyShitpost: Anything and everything goes.
RISA: Star Trek memes and shitposts

Community rules

Follow the site-wide rules and code of conduct
Be civil
Post Linux-related content
No recent reposts

Please report posts and comments that break these rules!

founded 1 year ago

MODERATORS

poopsmith

zephyr

338

Parsing HTML with regex (lemmy.sdf.org)

submitted 6 months ago by [email protected] to c/linuxmemes

40 comments fedilink hide all child comments

cross-posted from: https://lemmy.sdf.org/post/12950329

you are viewing a single comment's thread
view the rest of the comments

[–] [email protected] 21 points 6 months ago (13 children)

Putting a statement into old calligraphy is nice. But it all boils down to because I said so. If you're going to go to that effort you might as well put the rationale for why it can't possibly parse the language into the explanation rather than because I said so

[–] [email protected] 14 points 6 months ago

The text does technically give the reason on the first page:

It is not a regular language and hence cannot be parsed by regular expressions.

Here, "regular language" is a technical term, and the statement is correct.

The text goes on to discuss Perl regexes, which I think are able to parse at least all languages in LL(*). I'm fairly sure that is sufficient to recognize XML, but am not quite certain about HTML5. The WHATWG standard doesn't define HTML5 syntax with a grammar, but with a stateful parsing procedure which defies normal placement in the Chomsky hierarchy.

This, of course, is the real reason: even if such a regex is technically possible with some regex engines, creating it is extremely exhausting and each time you look into the spec to understand an edge case you suffer 1D6 SAN damage.

load more comments (12 replies)