*Quick Unicode quiz* :thinking_face: `allEqual()`...
# stdlib
d
Quick Unicode quiz 🤔
allEqual()
returns
true
if all elements are equal, and
allDistinct()
returns
true
if all elements are different. They exist for collections, sequences and arrays since Kotlin 2.4.20 (experimental), and we are thinking about adding them to
CharSequence
. Please answer without running the code. We want to know what you expect, not what Kotlin does today. Vote with reactions under each question, and please write in the thread of this message why you think so 🙏
y
I vaguely know that some emojis might take up multipe characters, especially when converted to UTF-16 1.
true
. If it fits in 16-bits, then obviously they're distinct, and if it doesn't, then surely the bit patterns aren't identical (something something surrogate pairs? Surely surrogate pairs consist of 2 different bit patterns, right?) 2.
true
. Just a hunch that 😲 might be a very early emoji, and thus it got some of the precious UTF-16 space. 3.
false
. IIRC, UC, trying to stay neutral geopolitically, decided to have flags as a combo of a special flag modifier, then 2 letters representing the country code. Since the flag modifier is repeated, it's clearly not distinct. 4.
false
. Pretty sure skin tone is another unicode modifier, so it's a separate character, hence you'll have 👍 skin tone 4::skin tone 4 or something like that. 5.
7
. I initially thought
5
, reasoning that it's the 4 family members and some modifier, but it's not an option. Then I vaguely remembered something about a "glue" character that glues these mega-emojis together, so it makes sense that you'd have 3 of that glue between the 4 family members
thank you color 1
h
I think I misunderstood the question 🤦‍♂️ I voted how I want things to be, not how I expect them to be currently...
o
1. they are distinct, unless these two specific emojis happen to share one of the surrogates 2. AFAIK all emojis are outside BMP, so every emoji is 2 code units (i. e. a surrogate pair). Each surrogate has a different bit prefix. The content of that string is
ABAB
- not all equal 3. Unless I am mistaken, flag emojis are a combination of
flag (universal symbol)
+
country
. Not distinct 4. Same as 2, but each of those is even more than 2 surrogates (like the simple emojis above). These are
thumbsup
+
color
-- 3 code units each? But still, not all equal. 5. That has to be like 11 code units long, no? Isn't it 4 emojis combined with a combinator character?
j
I think perhaps the functions shouldn't exist on
String
at all because they're almost certainly going to be used without understanding that characters are not grapheme clusters.
➕ 7
And if you make them operate on grapheme clusters, that means they somehow behave differently than every other string function.
d
Thanks @jw! I agree that grapheme clusters would make these functions different from every other string function, and Kotlin doesn't have any API for them yet anyway (we are working on this). So on
CharSequence
they can only work with
Chars
, and yes, all four examples above return
false
. You're right that people will often use them without knowing this. But people already write same checks by hand, for example
s.toSet().size == s.length
or
s.all { it == s[0] }
, and those have exactly the same problem. With a stdlib function it's at least documented. The KDoc says it compares Chars and the samples show the emoji cases. It's also faster than the `toSet()`/`asSequence()` workarounds (
allDistinct()
stops at the first duplicate,
allEqual()
doesn't allocate anything), and it's consistent with existing
Sequence<Char>
,
CharArray
,
List<Char>
,
Array<Char>
`allEqual`/`allDistinct` variants. WDYT this is enough to add them?
y
I've just realized I read
allDistinct
as
!allEqual
🤦🏼‍♂️ Not sure if it's likely for others to make the same mistake though I think in isolation
allDistinct
makes perfect sense. I just processed it wrong when contrasted with
allEqual
j
Is there a world where you'd suffix it like
allDistinctByChar()
maybe? I just worry it's a pretty big foot gun on string.
➕ 3
d
suffix it like
allDistinctByChar()
maybe?
Such an option was proposed. But, if these functions are on
CharSequence
, I expect them to compare UTF-16 Chars. It is even in the name, *Char*Sequence. And
CharSequence
already has a KDoc, that it manipulates with UTF-16 Chars.
💯 1
c
I think the examples results speak for themselves that
allEqual
/`allDistinct` are at the very least not behaving as people expect, so I think they would cause quite a few bugs that would go undetected for a long time. If possible, especially since you're working on grapheme clusters, it's probably a better idea to wait until that's released and provide both APIs
j
No one really interacts with CharSequence though, and the fact that it's an extension on it will be lost on most.
➕ 1
e
Java has
String.chars()
and
String.codePoints()
which return `Stream`s. if Kotlin had similar functions (presumably returning
Sequence
) then
.allDistinct()
etc. on them would be less surprising than on `String`/`CharSequence`
➕ 2