C++ float-to-int conversion can be undefined behavior (kttnr.net)

38 points by signa11 4 days ago

38 comments:

by digitalPhonix 4 days ago

Herb Sutter's comment on why it's ok is confusing to me:

> Regarding the use of UB internally: It's okay and if anyone is worried about it the use of UB is benign on the platforms we target (e.g., they don't involve hitting any hardware trap representations for these types)

Isn't the outcome of the UB (ie. whether it will "rm -rf /" or something else) dependent on both the target and the compiler? And the compiler (or future compiler) could plausibly make the assumption that the narrowing to an unrepresentable value will never occur and change behaviour because of it?

by wavemode 4 hours ago

For what it's worth, a GSL developer later reopened that GitHub issue and stated that they're going to look into fixing the UB. Sutter may have just been stating an assumption.

https://github.com/microsoft/GSL/issues/786#issuecomment-513...

> I'll raise this issue in the next internal GSL sync. I'd agree with y'all that this behavior: https://godbolt.org/z/4Tr1fe9xG is undesirable

by mort96 41 minutes ago

But ... surely Sutter ought to know better than to say "because the hardware handles this conversion reasonably, it's a benign case of UB"? Surely he knows that compilers can and will optimize based on the assumption that UB never happens?

The problem isn't, "oh no what if my CPU's float->int conversion instruction traps", that's an extremely naive way to think about UB. Everyone who has thought seriously about UB in C++ for any length of time knows this. It's worrying that this was Sutter's response.

by 20k 2 hours ago

Yeah Herb's 100% wrong here. Its common when people are downplaying the memory safety issues with C++ that they say things like this, but its completely incorrect. All invoked UB is potentially equally serious, and this is exploitable memory unsafety. Compilers can and do optimise away this kind of stuff (as other people have explained here)

There's also important context in that Herb is currently one of the people leading the current memory safety approach for C++

by jcranmer 3 hours ago

In LLVM, the result of floating-to-int conversion that is out of range of the int is a poison value, which means you get essentially the full unpredictability of UB.

That said, I'm a little hard-pressed to think of optimizations that would actually take advantage of poison, because floating-point range isn't really computed in the optimizer.

by mort96 27 minutes ago

Here's (my modified version of) an example someone came up with on lobste.rs: https://godbolt.org/z/e69b4Tqbs

I don't know exactly which optimization passes do what, but a few observations:

* The 'foo(unsigned int n)' function should never return a value that's greater than 'n', since it returns 'i < n ? i : n'.

* The value printed by the 'foo' function should always be the same as the value that's returned.

Yet the value it prints is 2700624104 (which is greater than 'n', which is 10 in this case), and the returned value is 2700623376, which is different. (The exact numbers vary run to run)

If the conversion "just" resulted in a bogus value, we would have expected some number <=10 to be printed two times.

by Maxatar 4 hours ago

Yes, this is all true but Sutter's comment is that the specific platforms that this specific implementation of the GSL targets results in the correct output. The platforms officially supported are:

GCC 12, 13, 14

XCode 14.3.1, 15.4

Clang 16, 17, 18

Visual Studio with MSVC VS2019, VS2022

Visual Studio with LLVM VS2019, VS2022

by afdbcreid 2 minutes ago

Then this is, unfortunately, entirely wrong. Here's an example causing a segfault when there is a bound checks that the compiler omits:

https://godbolt.org/z/8f6rv4dja

The example is adapted from a Rust example shown by @RalfJung in https://lobste.rs/s/ba2yfy/c_float_int_conversion_can_be_und....

by pjmlp 4 hours ago

No, UB is allowed special powers for compiler and standard library implementors, which is what Herb Sutter means with internal behaviour.

Meaning MSVC is aware of these cases, so the compiler has special cases for it.

by digitalPhonix 4 hours ago

That's my point - GSL is NOT MSVC only, it's a general purpose library and NOT a standard library implementation of a toolchain so any compiler is expected to be able to compile it (it also explicitly targets clang & gcc).

by pjmlp 4 hours ago

It was originally created by Microsoft and as Herb mentions "all our target platforms", so most likely it gets special treatment.

by aw1621107 3 hours ago

At the time Herb made that comment GSL was documented as supporting XCode 12.5.1/13.2.1, GCC 10/11, Clang 11/12, and Visual Studio 2019/2022 using both MSVC/LLVM [0]. Even if MSVC had special support for GSL I'm a bit more skeptical that such support would extend to XCode, GCC, and Clang.

[0]: https://github.com/microsoft/GSL/tree/99a29ce797c8337b8923f2...

by wavemode 4 hours ago

GSL is not the standard library nor an internal runtime library. Its GitHub page claims that it supports a variety of compilers:

> The GSL officially supports recent major versions of Visual Studio with both MSVC and LLVM, GCC, Clang, and XCode with Apple-Clang

by pjmlp 4 hours ago

It was originally created by Microsoft and as Herb mentions "all our target platforms", so most likely it gets special treatment.

by Maxatar 3 hours ago

There is absolutely no special treatment afforded to the GSL by any of the target platforms. While we can't inspect MSVC's source code, both clang and GCC do not have any support or affordance for the GSL whatsoever and it would be very unusual to expect MSVC's source code to have some kind of affordance for this library.

by achierius 3 hours ago

The 'special treatment' isn't technical, it's procedural -- insofar as if, during development for a new release, MSVC were to land some changes that broke GSL, Microsoft's testing would catch that and ensure that the changes were reverted or fixed to support the latter, prior to shipping. Since they're built as part of the same operating system, they can make sure not to step on one another's toes -- which is not a guarantee that they can make to third-party applications.

by tsimionescu 2 hours ago

If you've ever read anything about the internal culture in Microsoft, you'd know this is extremely implausible.

by Maxatar 3 hours ago

I have no idea where you possibly got this idea from since the Github Issues tracker for GSL has numerous instances of new releases of MSVC breaking GSL compilation.

by 20k 2 hours ago

Clang is a target for the GSL though. How can MSVC's special powers prevent this from being exploitable UB in Clang/LLVM?

This code boils down to static_cast<int>(some_double); so nothing fancy is going on here

by pjmlp 2 hours ago

Actually Microsoft has their own fork that ships with Visual Studio installer.

However that was me guessing from Herb Sutter's reply.

by LoganDark 4 hours ago

UB is bad not because it actually leads to any particular result on any particular platform or compiler, but because semantically it invalidates assumptions about a program. Rust is explicit on this, but it absolutely still applies to C/C++.

by em3rgent0rdr 2 hours ago

Well because it could lead to any result on some platform or compiler, it invalidates assumptions about the program.

by LoganDark 6 minutes ago

UB invalidates assumptions about the program not necessarily because it leads to arbitrary behavior in practice, but because it leads to arbitrary behavior in spirit. Even if there is no compiler in existence where the UB causes a problem, that does not make the program correct. UB is about whether the program is correctly defined, not about whether it works or not.

by gpvos 2 hours ago

Sounds like the standard should say that it results in an implementation-defined value (or wording to that effect). Saying it's UB gives the compilers way too much leeway.

by 20k 2 hours ago

Its incredibly hard to get changes like this into the standard, because there's a core contingent of people who seem to feel that UB is part of C++'s identity, and then there's very vague hand waving about performance. There is luckily a pretty successful push in wg21 to start removing a lot of the more unnecessary UB, so hopefully this gets sent to the sausage factory as well

by pjmlp 4 hours ago

Hopefully this will be part of UB fixes for C++29, where plenty of UB is being redefined as erroneous behaviour instead.

by dmitrygr 2 hours ago

> The correct fix is to bounds check before casting.

This will do wonders for speed. Actually explicitly using the safe isntr might be better. Something like this will happily compile to a single instr and cause you no grief even if the compiler had it out for you with UB. These instrs all clearly define outputs for all inputs (note that said outputs may not match across architectures)

   static inline __attribute__((always_inline)) int f2i(float myFloat) {
      int myInt;

      #if defined(__arm__)
         asm("VCVT.S32.F32 %0, %1":"=r"(myInt), "t"(myFloat));
      #elif defined (__aarch64__)
         asm("FCVTZS %0, %1":"=r"(myInt), "w"(myFloat));
      #elif defined (__x86_64__)
         asm("CVTTSS2SI %0, %1":"=r"(myInt), "x"(myFloat));
      #else
         #if 0 // be boring
            if (myFloat <= TOO_SMALL_FLOAT || myFloat => TOO_BIG_FLOAT)
               abort();
         #else
            #warning "Embrace the UB"
         #endif
         myInt = (int)myFloat;
      #endif
      return myInt;
   }
by orangepanda 4 hours ago

How could it be defined behaviour, when the result is different on ARM and x86?

by marcosdumay 4 hours ago

Architecture dependent is not the same as undefined.

by stouset 4 hours ago

Also, the spec says it’s undefined. But compiler authors can always special-case their own compilers.

by 20k 2 hours ago

Clang/LLVM at least treats this as UB, not as implementation defined behaviour however according to an LLVM dev

by cataphract 4 hours ago

Undefined behavior is not the same as implementation-defined or unspecified behavior. A program with undefined behavior is by definition an incorrect program. But there are cases where the spec actually gives some margin to the implementation. Programs relying on the choices of the implementation may be correct, even if non-portable.

by Maxatar 3 hours ago

>A program with undefined behavior is by definition an incorrect program.

This is simply false and an oft repeated myth. Undefined behavior has a specific technical definition that is in the C++ standard [1] and there is absolutely no mention in that definition or the implication of that definition that undefined behavior necessarily results in an invalid or incorrect program.

The definition of undefined behavior, right from the standard itself is... and I quote... get ready for it...

"behavior for which this document imposes no requirements"

That's it, nothing more, nothing less.

The standard even goes out of its way to state the following:

"Permissible undefined behavior ranges from ignoring the situation completely with unpredictable results, *to behaving during translation or program execution in a documented manner* characteristic of the environment".

Behaving in a documented manner characteristic of an environment is a far cry from being by incorrect by definition.

[1] https://www.open-std.org/jtc1/sc22/wg21/docs/papers/2024/n49...

by 20k 2 hours ago

>This document imposes no requirements on the behavior of programs that contain undefined behavior

That's saying that programs that exhibit undefined behaviour are not governed by the C++ spec. For a program to be a valid, spec governed piece of C++ code it has to exhibit no undefined behaviour (outside of some constraints). Its accurate to say that any undefined behaviour results in the code being executed no longer being C++, and it can have any behaviour. That's synonymous in common developer speak with 'incorrect', as its desirable for your C++ code to be executed as C++

by Joker_vD an hour ago

Do I love selective quoting! "Undefined behavior may be expected... when a program uses an incorrect construct or invalid data". There is also "erroneous behaviour" which "is always the consequence of incorrect program code".

Honestly, you'd have a better argument by quoting that "Correct execution" can include undefined behavior and erroneous behavior, depending on the data being processed". Which is quite a wild sentence to read, but here we are.

by cataphract 2 hours ago

I appreciate the correction. Although it should be said that modern implementations tend to opt to assume that UB never happens.

by lionkor 4 days ago

The core guidelines library is definitely not doing the right thing here. Very odd.

by functionmouse 2 hours ago

float considered harmful

Data from: Hacker News, provided by Hacker News (unofficial) API