beenull/ladybird

mirror of https://github.com/LadybirdBrowser/ladybird.git synced 2024-11-25 17:10:23 +00:00

Author	SHA1	Message	Date
Sam Atkins	b876e97719	AK: Expose the current position of a Utf8CodePointIterator as a pointer	2023-03-22 19:45:40 +01:00
Sam Atkins	067d0689c5	AK: Replace C-style casts	2023-03-09 21:43:54 +01:00
Timothy Flynn	434ca78425	AK: Protect Utf8View against inclusion in the Kernel It will soon be included in the Kernel by way of String.h. Utf8View includes DeprecatedString, which is not allowed in the Kernel.	2023-03-03 11:46:42 -05:00
Timothy Flynn	c4d78c29a2	AK: Invalidate overlong UTF-8 code point encodings For example, the code point U+002F could be encoded as UTF-8 with the bytes 0x80 0xAF. This trick has historically been used to bypass security checks.	2023-03-03 11:46:42 -05:00
Timothy Flynn	796a615bc1	AK: Replace UTF-8 string validation with a constexpr implementation This will allow validating UTF-8 strings at compile time, such as from String::from_utf8_short_string.	2023-03-03 11:46:42 -05:00
Timothy Flynn	0f20586346	AK: Add formatters for Utf8View and Utf32View Useful for debugging, especially in templated contexts.	2023-02-22 10:14:36 +01:00
Andreas Kling	6b497b8710	AK: Add two helpers to DeprecatedStringCodePointIterator	2023-01-29 23:41:42 +01:00
Andreas Kling	2dc657c77e	AK: Add DeprecatedStringCodePointIterator This is a safe iterator over the underlying code points. It will be used in Jakt to assist in the migration away from DeprecatedString.	2023-01-28 09:50:52 +01:00
Linus Groh	6e19ab2bbc	AK+Everywhere: Rename String to DeprecatedString We have a new, improved string type coming up in AK (OOM aware, no null state), and while it's going to use UTF-8, the name UTF8String is a mouthful - so let's free up the String name by renaming the existing class. Making the old one have an annoying name will hopefully also help with quick adoption :^)	2022-12-06 08:54:33 +01:00
Andreas Kling	ae3ffdd521	AK: Make it possible to not `using` AK classes into the global namespace This patch adds the `USING_AK_GLOBALLY` macro which is enabled by default, but can be overridden by build flags. This is a step towards integrating Jakt and AK types.	2022-11-26 15:51:34 +01:00
Andreas Kling	e7ba03ddd1	AK: Add Utf8View::iterator_at_byte_offset_without_validation() Unlike iterator_at_byte_offset(), this function assumes the provided byte offset is a valid offset into the UTF-8 character stream. This avoids walking the stream from the start.	2022-11-24 16:06:20 +00:00
Idan Horowitz	086969277e	Everywhere: Run clang-format	2022-04-01 21:24:45 +01:00
Idan Horowitz	4774bed589	AK: Make Utf8View constexpr-constructible	2022-01-17 14:46:07 +00:00
Andreas Kling	8b1108e485	Everywhere: Pass AK::StringView by value	2021-11-11 01:27:46 +01:00
Andreas Kling	391352c112	AK: Inline all the trivial Utf8View functions This improves parsing time on a large chunk of JS by ~3%.	2021-09-18 19:54:24 +02:00
Andreas Kling	1be4cbd639	AK: Make Utf8View constructors inline and remove C string constructor Using StringView instead of C strings is basically always preferable. The only reason to use a C string is because you are calling a C API.	2021-09-18 19:54:24 +02:00
Idan Horowitz	e8f6840471	AK+LibRegex: Disable construction of views from temporary Strings	2021-09-04 21:01:15 +02:00
Timothy Flynn	c4ee576531	AK: Add Utf8View::byte_offset_of overload for code point index lookups	2021-08-18 09:47:09 +04:30
Ali Mohammad Pur	55fa51b4e2	AK: Add a is_null() method to Utf{8,32}View Both of these can be null as well as empty, and there's a difference.	2021-07-18 21:10:55 +04:30
Idan Horowitz	fea6d952a4	AK: Add the Utf8View::{contains, trim} helper methods	2021-06-16 20:05:18 +01:00
DexesTTP	e01f1c949f	AK: Do not VERIFY on invalid code point bytes in UTF8View The previous behavior was to always VERIFY that the UTF-8 bytes were valid when iterating over the code points of an UTF8View. This change makes it so we instead output the 0xFFFD 'REPLACEMENT CHARACTER' code point when encountering invalid bytes, and keep iterating the view after skipping one byte. Leaving the decision to the consumer would break symmetry with the UTF32View API, which would in turn require heavy refactoring and/or code duplication in generic code such as the one found in Gfx::Painter and the Shell. To make it easier for the consumers to detect the original bytes, we provide a new method on the iterator that returns a Span over the data that has been decoded. This method is immediately used in the TextNode::compute_text_for_rendering method, which previously did this in a ad-hoc waay. This also add tests for the new behavior in TestUtf8.cpp, as well as reinforcements to the existing tests to check if the underlying bytes match up with their expected values.	2021-06-03 18:28:27 +04:30
Andreas Kling	12a42edd13	Everywhere: codepoint => code point	2021-06-01 10:01:11 +02:00
Andreas Kling	407d6cd9e4	AK: Rename Utf8CodepointIterator => Utf8CodePointIterator	2021-06-01 09:45:52 +02:00
Max Wipfli	14506e8f5e	AK: Implement Utf8CodepointIterator::peek(size_t) This adds a peek method for Utf8CodepointIterator, which enables it to be used in some parsing cases where peeking is necessary. peek(0) is equivalent to operator*, expect that peek() does not contain any assertions and will just return an empty Optional<u32>. This also implements a test case for iterating UTF-8.	2021-06-01 09:28:05 +02:00
Max Wipfli	a72bb34970	AK: Add Utf8View::iterator_at_byte_offset method This implements a method to get a Utf8CodepointIterator at a specified byte offset.	2021-05-21 21:57:03 +02:00
Max Wipfli	c1b452f754	AK: Add substring methods to Utf8View This patch implements a Unicode-safe substring method, which can be used when offset and length should be specified in actual characters instead of bytes. This can be used to mitigate issues where a string is split in the middle of a UTF-8 multi-byte character, which leads to invalid UTF-8. Furthermore, it implements to common shorthands for substring methods which take only an offset and return the substring until the end of the string.	2021-05-21 21:57:03 +02:00
Max Wipfli	8c19c2f296	AK: Change some argument and return types in Utf8View from int to size_t This changes the return type of Utf8View::byte_length and the argument types of substring_view from int to size_t.	2021-05-21 21:57:03 +02:00
Brian Gianforcaro	1682f0b760	Everything: Move to SPDX license identifiers in all files. SPDX License Identifiers are a more compact / standardized way of representing file license information. See: https://spdx.dev/resources/use/#identifiers This was done with the `ambr` search and replace tool. ambr --no-parent-ignore --key-from-file --rep-from-file key.txt rep.txt *	2021-04-22 11:22:27 +02:00
Idan Horowitz	edecf8f6a3	AK: Add starts_with to Utf8View Unlike String/StringView::starts_with this compares utf8 code points instead of "characters" (bytes), which is important when handling aribtary utf-8 input that could include overlong characters.	2021-03-25 10:59:34 +01:00
Lenny Maiorani	e6f907a155	AK: Simplify constructors and conversions from nullptr_t Problem: - Many constructors are defined as `{}` rather than using the ` = default` compiler-provided constructor. - Some types provide an implicit conversion operator from `nullptr_t` instead of requiring the caller to default construct. This violates the C++ Core Guidelines suggestion to declare single-argument constructors explicit (https://isocpp.github.io/CppCoreGuidelines/CppCoreGuidelines#c46-by-default-declare-single-argument-constructors-explicit). Solution: - Change default constructors to use the compiler-provided default constructor. - Remove implicit conversion operators from `nullptr_t` and change usage to enforce type consistency without conversion.	2021-01-12 09:11:45 +01:00
asynts	2927656d85	AK: Use size_t in methods of Utf8View.	2021-01-02 01:37:22 +01:00
Andreas Kling	13594b7146	LibGfx+AK: Make text elision work with multi-byte characters This was causing WindowServer and Taskbar to crash sometimes when the stars aligned and we tried cutting off a string ending with "..." right on top of an emoji. :^)	2020-12-28 23:54:10 +01:00
Lenny Maiorani	f5ced347e6	AK: Prefer using instead of typedef Problem: - `typedef` is a keyword which comes from C and carries with it old syntax that is hard to read. - Creating type aliases with the `using` keyword allows for easier future maintenance because it supports template syntax. - There is inconsistent use of `typedef` vs `using`. Solution: - Use `clang-tidy`'s checker called `modernize-use-using` to update the syntax to use the newer syntax. - Remove unused functions to make `clang-tidy` happy. - This results in consistency within the codebase.	2020-11-12 10:19:04 +01:00
Tom	6413acd78c	AK: Make Utf8View and Utf32View more consistent This enables use of these classes in templated code.	2020-10-22 15:23:45 +02:00
Nico Weber	ce95628b7f	Unicode: Try s/codepoint/code_point/g again This time, without trailing 's'. Ran: git grep -l 'codepoint' \| xargs sed -ie 's/codepoint/code_point/g	2020-08-05 22:33:42 +02:00
Nico Weber	19ac1f6368	Revert "Unicode: s/codepoint/code_point/g" This reverts commit `ea9ac3155d`. It replaced "codepoint" with "code_points", not "code_point".	2020-08-05 22:33:42 +02:00
Andreas Kling	ea9ac3155d	Unicode: s/codepoint/code_point/g Unicode calls them "code points" so let's follow their style.	2020-08-03 19:06:41 +02:00
Matthew Olsson	c831fb17bf	LibJS: Add StringIterator	2020-07-13 15:07:29 +02:00
Andreas Kling	23dad305e9	AK: Allow default-constructing Utf8View and Utf8CodepointIterator	2020-06-04 21:12:17 +02:00
AnotherTest	a4e0b585fe	AK: Add a way to get the number of valid bytes in a Utf8View	2020-05-18 11:31:43 +02:00
Andreas Kling	6cde7e4d20	AK: Add Utf8View::length_in_codepoints()	2020-05-17 13:05:39 +02:00
Sergey Bugaev	c0b32f7b76	Meta: Claim copyright for files created by me This changes copyright holder to myself for the source code files that I've created or have (almost) completely rewritten. Not included are the files that were significantly changed by others even though it was me who originally created them (think HtmlView), or the many other files I've contributed code to.	2020-01-24 15:15:16 +01:00
Andreas Kling	94ca55cefd	Meta: Add license header to source files As suggested by Joshua, this commit adds the 2-clause BSD license as a comment block to the top of every source file. For the first pass, I've just added myself for simplicity. I encourage everyone to add themselves as copyright holders of any file they've added or modified in some significant way. If I've added myself in error somewhere, feel free to replace it with the appropriate copyright holder instead. Going forward, all new source files should include a license header.	2020-01-18 09:45:54 +01:00
Andreas Kling	f4e6dae6fe	UTF-8: Add Utf8CodepointIterator::codepoint_length_in_bytes() This allows you to retrieve the length (in bytes) of the codepoint the iterator is currently pointing at.	2019-10-18 22:49:23 +02:00
Andreas Kling	fb39e46d3d	Utf8View: Try fixing the travis-ci build There's some overload ambiguity when doing Utf8View("literal")	2019-09-05 19:06:39 +02:00
Sergey Bugaev	c379f43d2a	AK: Add some more utility methods to Utf8View	2019-09-05 16:37:39 +02:00
Sergey Bugaev	5d3696174b	AK: Add a Utf8View type for iterating over UTF-8 codepoints Utf8View wraps a StringView and implements begin() and end() that return a Utf8CodepointIterator, which parses UTF-8-encoded Unicode codepoints and returns them as 32-bit integers. This is the first step towards supporting emojis in Serenity ^) https://github.com/SerenityOS/serenity/issues/490	2019-08-28 13:46:02 +02:00

47 commits