beenull/ladybird

mirror of https://github.com/LadybirdBrowser/ladybird.git synced 2024-11-22 15:40:19 +00:00

Author	SHA1	Message	Date
Timothy Flynn	325eabc770	LibRegex: Ensure the GoBack operation decrements the code unit index This was missed in commit `27d555bab0`.	2021-08-18 09:47:09 +04:30
Timothy Flynn	a9716ad44e	LibRegex: In non-Unicode mode, parse \u{4} as a repetition pattern	2021-08-18 09:47:09 +04:30
davidot	7613c22b06	LibJS: Add a mode to parse JS as a module In a module strict mode should be enabled at the start of parsing and we allow import and export statements.	2021-08-15 23:51:47 +01:00
Timothy Flynn	9509433e25	LibRegex: Implement and use a REPEAT operation for bytecode repetition Currently, when we need to repeat an instruction N times, we simply add that instruction N times in a for-loop. This doesn't scale well with extremely large values of N, and ECMA-262 allows up to N = 2^53 - 1. Instead, add a new REPEAT bytecode operation to defer this loop from the parser to the runtime executor. This allows the parser to complete sans any loops (for this instruction), and allows the executor to bail early if the repeated bytecode fails. Note: The templated ByteCode methods are to allow the Posix parsers to continue using u32 because they are limited to N = 2^20.	2021-08-15 11:43:45 +01:00
Timothy Flynn	f1ce998d73	LibRegex+LibJS: Combine named and unnamed capture groups in MatchState Combining these into one list helps reduce the size of MatchState, and as a result, reduces the amount of memory consumed during execution of very large regex matches. Doing this also allows us to remove a few regex byte code instructions: ClearNamedCaptureGroup, SaveLeftNamedCaptureGroup, and NamedReference. Named groups now behave the same as unnamed groups for these operations. Note that SaveRightNamedCaptureGroup still exists to cache the matched group name. This also removes the recursion level from the MatchState, as it can exist as a local variable in Matcher::execute instead.	2021-08-15 11:43:45 +01:00
Timothy Flynn	1a173be29d	LibRegex: Disallow unescaped quantifiers in Unicode mode	2021-08-15 11:43:45 +01:00
Timothy Flynn	c3e1f1f687	LibRegex: Use correct source characters for Unicode identity escapes	2021-08-15 11:43:45 +01:00
Timothy Flynn	6a485f612f	LibRegex: Implement legacy octal escape parsing closer to the spec The grammar for the ECMA-262 CharacterEscape is: CharacterEscape[U, N] :: ControlEscape c ControlLetter 0 [lookahead ∉ DecimalDigit] HexEscapeSequence RegExpUnicodeEscapeSequence[?U] [~U]LegacyOctalEscapeSequence IdentityEscape[?U, ?N] It's important to parse the standalone "\0 [lookahead ∉ DecimalDigit]" before parsing LegacyOctalEscapeSequence. Otherwise, all standalone "\0" patterns are parsed as octal, which are disallowed in Unicode mode. Further, LegacyOctalEscapeSequence should also be parsed while parsing character classes.	2021-08-15 11:43:45 +01:00
Timothy Flynn	83ca8c7e38	LibRegex: Convert LibRegex tests to use StringView in place of C-strings A subsequent commit will add tests that require a string containing only "\0". As a C-string, this will be interpreted as the null terminator. To make the diff for that commit easier to grok, this commit converts all tests to use StringView without any other functional changes.	2021-08-15 11:43:45 +01:00
Timothy Flynn	0c8f2f5aca	LibRegex: Ensure escaped hexadecimals are exactly 2 digits in length	2021-08-15 11:43:45 +01:00
Timothy Flynn	2e4b6fd1ac	LibRegex: Ensure escaped code points are exactly 4 digits in length	2021-08-15 11:43:45 +01:00
Timothy Flynn	e887314472	LibRegex: Fix ECMA-262 parsing of invalid identity escapes * Only alphabetic (A-Z, a-z) characters may be escaped with \c. The loop currently parsing \c includes code points between the upper/lower case groups. * In Unicode mode, all invalid identity escapes should cause a parser error, even in browser-extended mode. * Avoid an infinite loop when parsing the pattern "\c" on its own.	2021-08-15 11:43:45 +01:00
Brian Gianforcaro	a2a5cb0f24	AK: Add Time::is_negative() to detect negative time values	2021-08-15 12:20:38 +02:00
Daniel Bertalan	0a36cea9dc	Tests: Re-enable UserspaceEmulator tests on the Clang build Now that problems that made UE crash have been fixed, this test should now pass.	2021-08-14 18:42:14 +02:00
Itamar	e57fdb63f8	Tests: Add regression tests for the LibCpp preprocessor Similarly to the LibCpp parser regression tests, these tests run the preprocessor on the .cpp test files under Userland/LibCpp/Tests/preprocessor, and compare the output with existing .txt ground truth files.	2021-08-14 12:40:55 +02:00
Timothy Flynn	df14d11a11	LibRegex: Disallow invalid interval qualifiers in Unicode mode Fixes all remaining 'built-ins/RegExp/property-escapes' test262 tests.	2021-08-11 13:11:01 +02:00
Timothy Flynn	1e91334008	LibUnicode: Handle edge-case script extensions, Common and Inherited These script extensions have some peculiar behavior in the Unicode spec. The UCD ScriptExtension file does not contain these scripts. Rather, it is implied the code points which have these scripts as an extension are the code points that both: 1. Have Common or Inherited as their primary script value 2. Do not have any other script value in their script extension lists Because these are not explictly listed in the UCD, we must manually form these script extensions.	2021-08-11 13:11:01 +02:00
Timothy Flynn	47bb350ebd	LibUnicode: Generate separate tables for scripts and script extensions Notice that unlike the note in populate_general_category_unions(), script extension do indeed have code point ranges which overlap. Thus, this commit adds code to handle that, and hooks it into the GC unions.	2021-08-11 13:11:01 +02:00
Timothy Flynn	5ac23d244d	LibUnicode: Generate separate tables for Unicode properties Similar to General Categories, this generates separate tables for the Property list.	2021-08-11 13:11:01 +02:00
Timothy Flynn	b06c104076	LibUnicode: Include Unassigned code points in the Other General Category Now that the generator parses unassigned General Category properties, it can include Unassigned (Cn) in the Other (C) category.	2021-08-11 13:11:01 +02:00
Timothy Flynn	7dce2bfe23	LibUnicode: Generate separate tables for General Category properties Previously, each code point's General Category was part of the generated UnicodeData structure. This ultimately presented two problems, one functional and one performance related: * Some General Categories are applied to unassigned code points, for example the Unassigned (Cn) category. Unassigned code points are strictly excluded from UnicodeData.txt, so by relying on that file, the generator is unable to handle these categories. * Lookups for General Categories are slower when searching through the large UnicodeData hash map. Even though lookups are O(1), the hash function turned out to be slower than binary searching through a category-specific table. So, now a table is generated for each General Category. When querying a code point for a category, a binary search is done on each code point range in that category's table to check if code point has that category. Further, General Categories are now parsed from the UCD file DerivedGeneralCategory.txt. This file is a normal "prop list" file and contains the categories for unassigned code points.	2021-08-11 13:11:01 +02:00
Mandar Kulkarni	aaf232f903	Tests: Add test for String::bijective_base_from()	2021-08-09 14:14:07 +04:30
Daniel Bertalan	146dcf4856	Tests: Disable UserspaceEmulator tests for Clang builds There seems to be more incorrect assumptions about Clang-built executables' memory layout than expected. These make the CI fail even though the system is functional in all other aspects. While this is being fixed, let's just disable tests for UserspaceEmulator.	2021-08-08 10:55:36 +02:00
Daniel Bertalan	7396e4aedc	LibDebug: Store 64-bit numbers in AttributeValue This helps us avoid weird truncation issues and fixes a bug on Clang builds where truncation while reading caused the DIE offsets following large LEB128 numbers to be incorrect. This removes the need for the separate `LongUnsignedNumber` type.	2021-08-08 10:55:36 +02:00
Daniel Bertalan	5f2f460cc8	Tests: Add Clang pragma for turning off optimizations Clang does not accept `GCC optimize("O0")`, so it fails to build the system with it.	2021-08-08 10:55:36 +02:00
Itamar	4673a517f6	LibCpp: Do lexing in the Preprocessor We now call Preprocessor::process_and_lex() and pass the result to the parser. Doing the lexing in the preprocessor will allow us to maintain the original position information of tokens after substituting definitions.	2021-08-07 21:24:11 +02:00
Lenny Maiorani	8e949c5c91	Tests: Remove unused variables for clang build Problem: - Clang will not build `Tests/LibTLS` due to unused variables. Solution: - Remove the unused variables.	2021-08-06 23:55:27 +02:00
TheFightingCatfish	4e8e1b7b3a	AK: Improve the parsing of data urls Improve the parsing of data urls in URLParser to bring it more up-to- spec. At the moment, we cannot parse the components of the MIME type since it is represented as a string, but the spec requires it to be parsed as a "MIME type record".	2021-08-06 10:45:17 +02:00
Timothy Flynn	484ccfadc3	LibRegex: Support property escapes of Unicode script extensions	2021-08-04 13:50:32 +01:00
Timothy Flynn	06088df729	LibRegex: Support property escapes of the Unicode script property Note that unlike binary properties and general categories, scripts must be specified in the non-binary (Script=Value) form.	2021-08-04 13:50:32 +01:00
Brian Gianforcaro	4df1657898	Tests: Add coverage for sys$alarm() success case	2021-08-03 18:44:01 +02:00
Brian Gianforcaro	ea401fb3c3	Tests: Add coverage for sys$alarm() canceling a stale timer This is a regression test to validate the functionality that was reported broken in #9071, where the kernel would spin attempting to cancel a stale timer.	2021-08-03 18:44:01 +02:00
Timothy Flynn	dc9f516339	LibRegex: Generate negated property escapes as a single instruction These were previously generated as two instructions, Compare [Inverse] and Compare [Property].	2021-08-02 21:02:09 +04:30
Timothy Flynn	4de4312827	LibRegex: Support property escapes of the form \p{Type=Value} Before now, only binary properties could be parsed. Non-binary props are of the form "Type=Value", where "Type" may be General_Category, Script, or Script_Extension (or their aliases). Of these, LibUnicode currently supports General_Category, so LibRegex can parse only that type.	2021-08-02 21:02:09 +04:30
Timothy Flynn	1e10d6d7ce	LibRegex: Support property escapes of Unicode General Categories This changes LibRegex to parse the property escape as a Variant of Unicode Property & General Category values. A byte code instruction is added to perform matching based on General Category values.	2021-08-02 21:02:09 +04:30
Ali Mohammad Pur	85d87cbcc8	LibRegex: Add some tests for Fork{Stay,Jump} performance Without the previous fixes, these will blow up the stack.	2021-08-02 17:22:50 +04:30
Brian Gianforcaro	d1644c26d6	Tests: Remove unused header includes	2021-08-01 08:10:16 +02:00
Brian Gianforcaro	c54ae3afd6	Tests: Fix AK/TestJSON.cpp by not relying on disk resources The following commit broke Tests/AK/TestJSON.cpp as it removed the file that the test loaded from disk to validate JSON parsing. commit `ad141a2286` Author: Andreas Kling <kling@serenityos.org> Date: Sat Jul 31 15:26:14 2021 +0200 Base: Remove "test.frm" from HackStudio test project Instead of restoring the file, lets just embed a bit of JSON in the test case to avoid using external resources, as they obviously are surprising and make the test less portable across environments.	2021-07-31 23:56:40 +02:00
Timothy Flynn	d485cf29d7	LibRegex+LibUnicode: Begin implementing Unicode property escapes This supports some binary property matching. It does not support any properties not yet parsed by LibUnicode, nor does it support value matching (such as Script_Extensions=Latin).	2021-07-30 21:26:31 +01:00
Andreas Kling	bccdc08487	Kernel: Unmapping a non-mapped region with munmap() should be a no-op Not a regression per se from `0fcb9efd86` since we were crashing before that which is obviously worse.	2021-07-30 13:16:55 +02:00
Brian Gianforcaro	c9395d7e9a	Tests: Validate unmapping 0x0 doesn't crash the Kernel Previously unmapping any offset starting at 0x0 would assert in the kernel, add a regression test to validate the fix. Co-authored-by: Federico Guerinoni <guerinoni.federico@gmail.com>	2021-07-30 11:28:55 +02:00
Timothy Flynn	c4bfda7f7f	LibUnicode: Handle code points that are both cased and case-ignorable Apparently, some code points fit both categories, for example U+0345 (COMBINING GREEK YPOGEGRAMMENI). Handle this fact when determining if a code point is a final code point in a string.	2021-07-28 23:42:29 +02:00
Timothy Flynn	7827aede6f	LibUnicode: Check word break when deciding on case-ignorable code points	2021-07-28 23:42:29 +02:00
Timothy Flynn	c45a014645	LibUnicode: Check property list when deciding if a code point is cased	2021-07-28 23:42:29 +02:00
ovf	898b8ffcb6	LibWeb: Avoid assertion failure on parsing numeric character references	2021-07-28 18:32:22 +02:00
Timothy Flynn	39f971e42b	LibUnicode: Begin implementing special Unicode case folding This implements unconditional special case folding, and conditional folding for non-locale cases. Worth noting that the only conditional, non-locale special case is for converting an uppercase sigma to lowercase.	2021-07-27 21:04:36 +01:00
ovf	13c7d55320	LibWeb: Fix parsing of character references in attribute values	2021-07-27 00:03:43 +02:00
Timothy Flynn	4dda3edc9e	LibUnicode: Introduce a Unicode library for interacting with UCD files The Unicode standard publishes the Unicode Character Database (UCD) with information about every code point, such as each code point's upper case mapping. LibUnicode exists to download and parse UCD files at build time and to provide accessors to that data. As a start, LibUnicode includes upper- and lower-case code point converters.	2021-07-26 17:03:55 +01:00
brapru	7e40c17460	AK: Create MACAddress from string Previously there was no way to create a MACAddress by passing a direct address as a string. This will allow programs like the arp utility to create a MACAddress instance by user-passed addresses.	2021-07-25 17:57:08 +02:00
Luke	a00b5fc7b7	Tests: Add tests for the quoted printable decoder	2021-07-24 20:11:28 +04:30

1 2 3 4 5

242 commits