Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions standard/attributes.md
Original file line number Diff line number Diff line change
@@ -1,9 +1,9 @@
# 23 Attributes

## 23.1 General

Check warning on line 4 in standard/attributes.md

View workflow job for this annotation

GitHub Actions / Markdown to Word Converter

standard/attributes.md#L4

MDC032::Line length 86 > maximum 81
Much of the C# language enables the programmer to specify declarative information about the entities defined in the program. For example, the accessibility of a method in a class is specified by decorating it with the *method_modifier*s `public`, `protected`, `internal`, and `private`.

Check warning on line 6 in standard/attributes.md

View workflow job for this annotation

GitHub Actions / Markdown to Word Converter

standard/attributes.md#L6

MDC032::Line length 92 > maximum 81
C# enables programmers to invent new kinds of declarative information, called ***attribute***s. Programmers can then attach attributes to various program entities, and retrieve attribute information in a run-time environment.

> *Note*: For instance, a framework might define a `HelpAttribute` attribute that can be placed on certain program elements (such as classes and methods) to provide a mapping from those program elements to their documentation. *end note*
Expand Down Expand Up @@ -826,7 +826,7 @@

The attribute `System.Runtime.CompilerServices.CallerFilePathAttribute` is allowed on optional parameters when there is a standard implicit conversion ([§10.4.2](conversions.md#1042-standard-implicit-conversions)) from `string` to the parameter’s type.

If a function invocation from a location in source code omits an optional parameter with the `CallerFilePathAttribute`, then a string literal representing that location’s file path is used as an argument to the invocation instead of the default parameter value.
If a function invocation from a location in source code omits an optional parameter with the `CallerFilePathAttribute`, then a UTF-16 string literal representing that location’s file path is used as an argument to the invocation instead of the default parameter value.

The format of the file path is implementation-dependent.

Expand All @@ -836,7 +836,7 @@

The attribute `System.Runtime.CompilerServices.CallerMemberNameAttribute` is allowed on optional parameters when there is a standard implicit conversion ([§10.4.2](conversions.md#1042-standard-implicit-conversions)) from `string` to the parameter’s type.

If a function invocation from a location within the body of a function member or within an attribute applied to the function member itself or its return type, parameters or type parameters in source code omits an optional parameter with the `CallerMemberNameAttribute`, then a string literal representing the name of that member is used as an argument to the invocation instead of the default parameter value.
If a function invocation from a location within the body of a function member or within an attribute applied to the function member itself or its return type, parameters or type parameters in source code omits an optional parameter with the `CallerMemberNameAttribute`, then a UTF-16 string literal representing the name of that member is used as an argument to the invocation instead of the default parameter value.

> *Note*: In the case of a function invocation from a top-level statement the string is a representation of the implementation provided name ([§7.1.3](basic-concepts.md#713-using-top-level-statements)). *end note*

Expand Down
34 changes: 28 additions & 6 deletions standard/expressions.md
Original file line number Diff line number Diff line change
@@ -1,15 +1,15 @@
# 12 Expressions

Check warning on line 1 in standard/expressions.md

View workflow job for this annotation

GitHub Actions / Markdown to Word Converter

standard/expressions.md#L1

MDC032::Line length 83 > maximum 81

Check warning on line 1 in standard/expressions.md

View workflow job for this annotation

GitHub Actions / Markdown to Word Converter

standard/expressions.md#L1

MDC032::Line length 84 > maximum 81

Check warning on line 2 in standard/expressions.md

View workflow job for this annotation

GitHub Actions / Markdown to Word Converter

standard/expressions.md#L2

MDC032::Line length 84 > maximum 81

Check warning on line 2 in standard/expressions.md

View workflow job for this annotation

GitHub Actions / Markdown to Word Converter

standard/expressions.md#L2

MDC032::Line length 85 > maximum 81
## 12.1 General

Check warning on line 3 in standard/expressions.md

View workflow job for this annotation

GitHub Actions / Markdown to Word Converter

standard/expressions.md#L3

MDC032::Line length 85 > maximum 81

Check warning on line 3 in standard/expressions.md

View workflow job for this annotation

GitHub Actions / Markdown to Word Converter

standard/expressions.md#L3

MDC032::Line length 86 > maximum 81

An expression is a sequence of operators and operands. This clause defines the syntax, order of evaluation of operands and operators, and meaning of expressions.

Check warning on line 5 in standard/expressions.md

View workflow job for this annotation

GitHub Actions / Markdown to Word Converter

standard/expressions.md#L5

MDC032::Line length 82 > maximum 81

Check warning on line 5 in standard/expressions.md

View workflow job for this annotation

GitHub Actions / Markdown to Word Converter

standard/expressions.md#L5

MDC032::Line length 102 > maximum 81

Check warning on line 6 in standard/expressions.md

View workflow job for this annotation

GitHub Actions / Markdown to Word Converter

standard/expressions.md#L6

MDC032::Line length 97 > maximum 81

Check warning on line 6 in standard/expressions.md

View workflow job for this annotation

GitHub Actions / Markdown to Word Converter

standard/expressions.md#L6

MDC032::Line length 82 > maximum 81

Check warning on line 6 in standard/expressions.md

View workflow job for this annotation

GitHub Actions / Markdown to Word Converter

standard/expressions.md#L6

MDC032::Line length 82 > maximum 81
An expression *E* is said to ***directly contain*** a subexpression *E₁* if it is not subject to a user-defined conversion [§10.5](conversions.md#105-user-defined-conversions) whose parameter is not of a non-nullable value type, and one of the following conditions holds:

Check warning on line 8 in standard/expressions.md

View workflow job for this annotation

GitHub Actions / Markdown to Word Converter

standard/expressions.md#L8

MDC032::Line length 90 > maximum 81
- *E* is *E₁*.
- If *E* is a parenthesized expression `(E₂)`, and *E₂* directly contains *E₁*.
- If *E* is a null-forgiving operator expression `E₂!`, and *E₂* directly contains *E₁*.
- If *E* is a cast expression `(T)E₂`, and the cast does not subject *E₂* to a non-lifted user-defined conversion whose parameter is not of a non-nullable value type, and *E₂* directly contains *E₁*.

Check warning on line 12 in standard/expressions.md

View workflow job for this annotation

GitHub Actions / Markdown to Word Converter

standard/expressions.md#L12

MDC032::Line length 82 > maximum 81

## 12.2 Expression classifications

Expand Down Expand Up @@ -4287,7 +4287,7 @@

For an operation of the form `x + y`, binary operator overload resolution ([§12.4.5](expressions.md#1245-binary-operator-overload-resolution)) is applied to select a specific operator implementation. The operands are converted to the parameter types of the selected operator, and the type of the result is the return type of the operator.

The predefined addition operators are listed below. For numeric and enumeration types, the predefined addition operators compute the sum of the two operands. When one or both operands are of type `string`, the predefined addition operators concatenate the string representation of the operands.
The predefined addition operators are listed below. For numeric and enumeration types, the predefined addition operators compute the sum of the two operands. When one or both operands are of type `string`, the predefined addition operators concatenate the string representation of the operands. When both operands are of type `ReadOnlySpan<byte>` and both are semantically UTF-8 byte representations, the predefined addition operator concatenates the bytes of the operands.

- Integer addition:

Expand Down Expand Up @@ -4337,16 +4337,16 @@
```

At run-time these operators are evaluated exactly as `(E)((U)x + (U)y`).
- String concatenation:
- UTF-16 string concatenation:

```csharp
string operator +(string x, string y);
string operator +(string x, object y);
string operator +(object x, string y);
```

These overloads of the binary `+` operator perform string concatenation. If an operand of string concatenation is `null`, an empty string is substituted. Otherwise, any non-`string` operand is converted to its string representation by invoking the virtual `ToString` method inherited from type `object`. If `ToString` returns `null`, an empty string is substituted.

These overloads of the binary `+` operator perform concatenation of UTF-16 strings. If an operand is `null`, an empty UTF-16 string is substituted. Otherwise, any non-`string` operand that is not a ref struct ([§16.2.3](structs.md#1623-ref-modifier)) is converted to its UTF-16 string representation by invoking the virtual `ToString` method inherited from type `object`. If `ToString` returns `null`, an empty UTF-16 string is substituted.
> *Example*:
>
> <!-- Example: {template:"standalone-console", name:"AdditionOperator", expectedOutput:["s = ><","i = 1","f = 1.23E+15","d = 2.900"]} -->
Expand Down Expand Up @@ -4374,7 +4374,29 @@
>
> *end example*

The result of the string concatenation operator is a `string` that consists of the characters of the left operand followed by the characters of the right operand. The string concatenation operator never returns a `null` value. A `System.OutOfMemoryException` may be thrown if there is not enough memory available to allocate the resulting string.
The result of the operator is a `string` that consists of the characters of the left operand followed by the characters of the right operand. The string concatenation operator never returns a `null` value. A `System.OutOfMemoryException` may be thrown if there is not enough memory available to allocate the resulting string.
- UTF-8 string concatenation:

```csharp
ReadOnlySpan<byte> operator +(ReadOnlySpan<byte> x, ReadOnlySpan<byte> y);
```

This overload of the binary `+` operator performs concatenation of UTF-8 string literals and the results of other applications of this operator (which is much more restrictive than UTF-16 string concatenation). It is applicable if and only if both operands are *semantically UTF-8 byte representations*. An operand is *semantically a UTF-8 byte representation* if it is a UTF-8 string literal, the result of an application of this operator, or a parenthesized expression whose enclosed expression is *semantically a UTF-8 byte representation*.

The result of the operator is a `ReadOnlySpan<byte>` that consists of the bytes of the left operand followed by the bytes of the right operand. The result is itself *semantically a UTF-8 byte representation*, and so may be used as an operand to a further application of this operator.

> *Example*:
>
> <!-- Example: {template:"standalone-console", name:"AdditionOperator2", expectedErrors:["CS9047","CS9047"]} -->
> ```csharp
> ReadOnlySpan<byte> sp1 = "ABC"u8 + "DEF"u8; // OK
> ReadOnlySpan<byte> sp2 = sp1 + "DEF"u8; // error
> ReadOnlySpan<byte> sp3 = "ABC"u8 + "DEF"u8 + "123"u8; // OK
> ReadOnlySpan<byte> sp4 = "ABC"u8 + (ReadOnlySpan<byte>)stackalloc byte[]
> { (byte)'D', (byte)'E', (byte)'F', (byte)'\x0' }; // error
> ```
>
> In the case of `sp1`, both operands are UTF-8 string literals. However, once `sp1` is initialized, that UTF-8 pedigree is no longer tracked. That is, `sp1` itself is not seen as being UTF-8 encoded. As such, it is not permitted to be an operand in the case of the initialization of `sp2`. In the initializer for `sp3`, the left pair of operands is evaluated, and as they are both UTF-8 string literals, the result is deemed to also be UTF-8 encoded, so it can further be used as the left operand of the right operator. In the case of `sp4`, while both operands are `ReadOnlySpan<byte>`s, only the left operand is UTF-8 encoded, even though the `Span<byte>` returned by `stackalloc` has the internal form of a UTF-8 string literal (that is, an array of bytes with a null-byte terminator). See [§6.4.5.6](lexical-structure.md#6456-string-literals). *end example*
- Delegate combination. Every delegate type implicitly provides the following predefined operator, where `D` is the delegate type:

```csharp
Expand Down Expand Up @@ -7516,7 +7538,7 @@

Only the following constructs are permitted in constant expressions:

- Literals (including the `null` literal).
- Literals (including the `null` literal, but excluding UTF-8 string literals).
- Constant interpolated strings.
- References to `const` members of class, struct, and interface types.
- References to members of enumeration types.
Expand Down
31 changes: 28 additions & 3 deletions standard/lexical-structure.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@
Conceptually speaking, a program is compiled using three steps:

1. Transformation, which converts a file from a particular character repertoire and encoding scheme into a sequence of Unicode characters.
1. Lexical analysis, which translates a stream of Unicode input characters into a stream of tokens.

Check warning on line 10 in standard/lexical-structure.md

View workflow job for this annotation

GitHub Actions / Markdown to Word Converter

standard/lexical-structure.md#L10

MDC032::Line length 85 > maximum 81
1. Syntactic analysis, which translates the stream of tokens into executable code.

Apart from accepting UTF-8 encoded input (as required by [§5](conformance.md#5-conformance), a conforming implementation may choose to accept and transform additional character encoding schemes (such as UTF-16, UTF-32, or non-Unicode character mappings).
Expand Down Expand Up @@ -920,14 +920,16 @@

In a verbatim string literal, the characters between the delimiters are interpreted verbatim, with the only exception being a *Quote_Escape_Sequence*, which represents one double-quote character. In particular, simple escape sequences, and hexadecimal and Unicode escape sequences are not processed in verbatim string literals. A verbatim string literal may span multiple lines.

All string literal forms may optionally have a trailing *Utf8_Suffix*. The representation of each form is discussed below.

```ANTLR
String_Literal
: Regular_String_Literal
| Verbatim_String_Literal
;

fragment Regular_String_Literal
: '"' Regular_String_Literal_Character* '"'
: '"' Regular_String_Literal_Character* '"' Utf8_Suffix?
;

fragment Regular_String_Literal_Character
Expand All @@ -943,7 +945,7 @@
;

fragment Verbatim_String_Literal
: '@"' Verbatim_String_Literal_Character* '"'
: '@"' Verbatim_String_Literal_Character* '"' Utf8_Suffix?
;

fragment Verbatim_String_Literal_Character
Expand All @@ -958,6 +960,10 @@
fragment Quote_Escape_Sequence
: '""'
;

fragment Utf8_Suffix
: 'u8' | 'U8'
;
```

> *Example*: The example
Expand Down Expand Up @@ -990,7 +996,26 @@
<!-- markdownlint-enable MD028 -->
> *Note*: Since a hexadecimal escape sequence can have a variable number of hex digits, the string literal `"\x123"` contains a single character with hex value `123`. To create a string containing the character with hex value `12` followed by the character `3`, one could write `"\x00123"` or `"\x12"` + `"3"` instead. *end note*

The type of a *String_Literal* is `string`.
A *String_Literal* that does not contain a *Utf8_Suffix* is a ***UTF-16 string literal***, whose type is `string`.

A *String_Literal* that contains a *Utf8_Suffix* is a ***UTF-8 string literal***, whose type is `System.ReadOnlySpan<byte>` (an indexable collection type), and whose value contains a UTF-8 byte representation of the string. A null terminator (a byte with value zero) is placed beyond the last byte in memory (and outside the length of the `ReadOnlySpan<byte>`) in order to support scenarios that expect null-terminated byte strings. A UTF-8 string literal is not a constant. A UTF-8 string literal without its *Utf8_Suffix* shall be valid UTF-16. (For example, `"\uDC00\uDD00"u8` is ill-formed as one low surrogate cannot be followed by another.)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should add a new subclause within 16.5.15 Safe context constraint, saying that when a UTF-8 string literal has a ReadOnlySpan<byte> value, the safe-context of that value is caller-context. That would allow it to be returned out of the function, like so:

using System;
public class C {
    public ReadOnlySpan<byte> M() {
        return "xyz"u8;
    }
}

Concatenation of UTF-8 string literals would then follow the safe-context rule in 16.5.15.5 Operators. As the safe-context of both operands would be caller-context, the safe-context of the result would be the same, without having to be specified separately for this case.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The following would then likewise be valid:

public class C {
    public ref readonly byte M() {
        return ref "xyz"u8[0];
    }
}


> *Note*: While every UTF-8 string literal is a `ReadOnlySpan<byte>`, not every `ReadOnlySpan<byte>` represents a UTF-8 string literal. See the description of UTF-8 string concatenation in [§12.13.5](expressions.md#12135-addition-operator). *end note*
<!-- markdownlint-disable MD028 -->

<!-- markdownlint-enable MD028 -->
> *Note*: Because `ReadOnlySpan<byte>` is a ref struct type, the value of a UTF-8 string literal cannot be implicitly converted to `object`, nor can `ReadOnlySpan<byte>` be used as a type argument ([§16.2.3](structs.md#1623-ref-modifier)). *end note*
<!-- markdownlint-disable MD028 -->

<!-- markdownlint-enable MD028 -->
> *Example*: Here are examples of each form of string literal:
>
> | **Encoding** | **Type** | **Regular String Literal** | **Verbatim String Literal** | **Raw String Literal** |
> |--------------|----------------------|---------------------|--------------------|--------------------|
> | UTF-16 | `string` | `"Hello"` | `@"Hello"` | `"""Hello"""` |
> | UTF-8 | `ReadOnlySpan<byte>` | `"Hello"u8` | `@"Hello"u8` | `"""Hello"""u8` |
>
> *end example*

Each string literal does not necessarily result in a new string instance. When two or more string literals that are equivalent according to the string equality operator ([§12.15.8](expressions.md#12158-string-equality-operators)), appear in the same assembly, these string literals refer to the same string instance.

Expand Down
2 changes: 2 additions & 0 deletions standard/structs.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# 16 Structs

## 16.1 General

Check warning on line 3 in standard/structs.md

View workflow job for this annotation

GitHub Actions / Markdown to Word Converter

standard/structs.md#L3

MDC032::Line length 82 > maximum 81

Structs are similar to classes in that they represent data structures that can contain data members and function members. However, unlike classes, structs are value types and do not require heap allocation. A variable of a `struct` type directly contains the data of the `struct`, whereas a variable of a class type contains a reference to the data, the latter known as an object.

Expand All @@ -11,8 +11,8 @@
## 16.2 Struct declarations

### 16.2.1 General

Check warning on line 14 in standard/structs.md

View workflow job for this annotation

GitHub Actions / Markdown to Word Converter

standard/structs.md#L14

MDC032::Line length 82 > maximum 81
A *struct_declaration* is a *type_declaration* ([§14.8](namespaces.md#148-type-declarations)) that declares a new struct:

Check warning on line 15 in standard/structs.md

View workflow job for this annotation

GitHub Actions / Markdown to Word Converter

standard/structs.md#L15

MDC032::Line length 91 > maximum 81

```ANTLR
struct_declaration
Expand Down Expand Up @@ -946,6 +946,8 @@

A `default` expression, for any type, has safe-context of caller-context.

A UTF-8 string literal ([§6.4.5.6](lexical-structure.md#6456-string-literals)) has a safe-context of caller-context.

For any non-default expression whose compile-time type is a ref struct has a safe-context defined by the following sections.

The safe-context records which context a value may be copied into. Given an assignment from an expression `E1` with a safe-context `S1`, to an expression `E2` with safe-context `S2`, it is an error if `S2` is a wider context than `S1`.
Expand Down
Loading