# Why no builtin string type?

**URL:** <https://ziggit.dev/t/why-no-builtin-string-type/5326>\
**Category:** Explain\
**Tags:** standard-library\
**Created:** [July 23, 2024, 10:54pm UTC](https://ziggit.dev/t/why-no-builtin-string-type/5326 "2024-07-23T22:54:07Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![squeek502](https://ziggit.dev/user_avatar/ziggit.dev/squeek502/32/409_2.png) [@squeek502](https://ziggit.dev/u/squeek502)\
**Post date:** [July 24, 2024, 12:00am UTC](https://ziggit.dev/t/why-no-builtin-string-type/5326/2 "2024-07-24T00:00:30Z")

</div>

Because strings are _complicated_.

That `zig-string` library you linked, for example, would mishandle comparison:

```zig
var myString = String.init(allocator);
defer myString.deinit();

try myString.concat("Ç");
assert(myString.cmp("Ç"));

```

This assertion would fail, even though the strings appear to be identical. That’s because the first uses [Normalization Form D](https://www.unicode.org/reports/tr15/): `C` ([U+0043](https://www.compart.com/en/unicode/U+0043)) + `◌̧` ([U+0327](https://www.compart.com/en/unicode/U+0327)), while the second uses Normalization Form C: `Ç` ([U+00C7](https://www.compart.com/en/unicode/U+00C7)). To actually compare UTF-8 strings in ways a human might expect, decisions about normalization need to be made.

The above is just one example. This series of articles by @dude_the_builder details the complication of Unicode well:

- [Unicode Basics in Zig - Zig NEWS](https://zig.news/dude_the_builder/unicode-basics-in-zig-dj3)
- [Ziglyph Unicode Wrangling - Zig NEWS](https://zig.news/dude_the_builder/ziglyph-unicode-wrangling-llj)
- [Unicode String Operations - Zig NEWS](https://zig.news/dude_the_builder/unicode-string-operations-536e)

(note that ziglyph has now been superseded by [zg](https://codeberg.org/dude_the_builder/zg))

So, for Zig to have a ‘proper’ UTF-8 String implementation, it would need to embed the [Unicode data](https://www.unicode.org/ucd/) and deal with all the complications of dealing with Unicode. My understanding is that’s not something that Zig-the-language or Zig-the-standard-library wants to take on if it doesn’t have to (especially since the Unicode data is a moving target).

Additionally, a UTF-8 String type is unable to handle arbitrary data, meaning the String type could not be used for a lot of the things Zig cares about: file paths, environment variables, etc. See [Fix handling of Windows (WTF-16) and WASI (UTF-8) paths, etc by squeek502 · Pull Request #19005 · ziglang/zig · GitHub](https://github.com/ziglang/zig/pull/19005) for more details on that sort of thing.

---

_[View the full topic](https://ziggit.dev/t/why-no-builtin-string-type/5326)._
