S tandard ECMA-121 2 n d Edition - December 2000
Standardizing Information
and
Communication
Systems
8-Bit Single-Byte Coded Graphic Character sets: Latin/Hebrew Alphabet
Phone: +41 22 849.60.00 - Fax: +41 22 849.60.01 - URL: http://www.ecma.ch - Internet: [email protected]
.
S tandard ECMA-121 2 n d Edition - December 2000
Standardizing
Information
and
Communication
Systems
8-Bit Single-Byte Coded Graphic Character sets: Latin/Hebrew Alphabet
Phone: +41 22 849.60.00 - Fax: +41 22 849.60.01 - URL: http://www.ecma.ch - Internet: [email protected] MB
ECMA-121.DOC
19-12-00 17,42
.
Brief History
The adoption of Standard ECMA-6 (ISO 646) in 1965 as the agreed international 7-bit code for information interchange has led to the development of many national, international and application-oriented versions of this code which have been in wide use for quite some time. These versions had a number of limitations generally inherent to the size of the code: −
they did not provide all graphic characters which may be needed,
−
for some characters, specially for accented letters, it was necessary to resort to BACKSPACE sequences, which created problems when processing data containing such composite characters,
−
interchange among different versions was practically limited to the 82 common graphic characters.
With the advent of 8-bit coding it was possible to increase the number of graphic characters. ISO 6937/2, for example, provided a character set covering the requirements of most languages based on the Latin alphabet. This character set, although well suited for text communication, was difficult to use for processing as some graphic characters were represented by one and others by two bit combinations. Thus, the need was recognized for coded graphic character sets, each of which: −
is the same for all users of a given area,
−
provides single-byte coding of all graphic characters thus permitting easy processing,
−
takes into account character sets used in the industry.
Since 1982 the urgency of the need for an 8-bit single-byte coded character set was recognized in ECMA as well as in ANSI/X3L2 and numerous working papers were exchanged between the two groups. In February 1984 ECMA TC1 submitted to ISO/TC97/SC2 (which has become ISO/IEC JTC 1/SC2 in 1987) a proposal for such a coded character set. At its meeting of April 1984 SC2 decided to propose a new item of work for this topic. Technical discussions during and after this meeting led TC1 to adopt the coding scheme proposed by X3L2. International Standard ISO/IEC 8859-1 is based on this joint ANSI/ECMA proposal. ECMA published its corresponding Standard ECMA-94 in March 1985. After this first publication, the work of ECMA TC1 on further coded graphic character sets has led to the following results: i.
The present Standard ECMA-121 for a Latin/Hebrew coded graphic set. This 2 nd Edition has been developed to keep it fully aligned with the new edition of ISO/IEC 8859-6.
ii. The second edition of Standard ECMA-94 comprising four coded graphic character sets for the Latin script, identified as Latin Alphabets No. 1 to No. 4. These alphabets have a number of characters in common, in particular those allocated to columns 02 to 07. These four Latin Alphabets have been submitted to ISO/IEC and JTC 1 and have become Parts 1 to 4 of ISO/IEC 8859. iii. A series of ECMA Standards for coded graphic character sets comprising those characters of the Latin Alphabets allocated to columns 02 to 07 and characters of another script for multiple-language applications. These ECMA Standards cover the Arabic, Cyrillic, and Greek scripts. These ECMA Standards ECMA-113, ECMA-114, and ECMA-118, resp., have become Parts 5 to 7, resp., of ISO/IEC 8859. iv. Latin Alphabets No. 5 and No. 6 have been published as ECMA-128 and ECMA-144, resp. They have become Parts 9 and 10, resp., of ISO/IEC 8859. This ECMA Standard has been adopted as 2 nd edition of Standard ECMA-121 by the ECMA General Assembly of December 2000.
- i -
Table of contents
1 2
Scope Conformance 2 . 1 Co n f o r ma n c e o f in f o r ma tio n in te r c h a n g e 2 . 2 Co n f o r ma n c e o f d e v ic e s 2.2.1 D e v ic e d e s c r ip tio n 2.2.2 O r ig in a tin g d e v ic e s 2.2.3 Re c e i v i n g d e v i c e s
1 1 1 1 1 1 1
3
References
1
4
Definitions b i- d ir e c tio n a l te x t b it co mb i n a t i o n b yte character c o d e ta b le coded character set; code c o d e d - c h a r a c t e r - d a t a - e l e me n t ( C C - d a t a - e l e me n t ) directional character properties graphic character g r a p h ic s ymb o l imp lic i t d ir e c t i o n a l i t y left-to-right character p o s itio n right-to-left character
2 2 2 2 2 2 2 2 2 2 2 2 3 3 3
4.1 4.2 4.3 4.4 4.5 4.6 4.7 4.8 4.9 4.10 4.11 4.12 4.13 4.14 5
Notation, code table and names 5 . 1 N o ta tio n 5 . 2 L a yo u t o f th e c o d e ta b le 5 . 3 N a me s a n d me a n in g s . 5.3.1 SPACE (SP) 5.3.2 NO-BREAK SPACE (NBSP) 5.3.3 SOFT HYPHEN (SHY) 5.3.4 LEFT-TO-RIGHT MARK (LRM) 5.3.5 RI G H T - T O - L E F T MA R K ( R L M)
6 6.1 6.2 7 7.1 7.2
Specification of the coded character set Ch a r a c t e r s o f t h e s e t a n d t h e i r c o d e d r e p r e s e n t a t i o n Co d e ta b le
3 3 3 3 4 4 4 4 4 4 4 8
Identification of the character set 9 Identification according to ECMA-35 and ECMA-43 9 I d e n tif ic a tio n u s in g th e I S O I n te r n a tio n a l r e g is te r o f c o d e d c h a r a c te r s e ts to b e u s e d with escape sequences 10
A n n e x A - C o v e r a g e o f la n g u a g e s
11
- ii -
Annex B .- Main differences between the first edition and this second edition of ECMA-121
13
Annex C - Bi-directional text support
15
Annex D - Bibliography
17
1
Scope This ECMA Standard specifies a set of 155 coded graphic characters identified as the Latin/Hebrew alphabet. This set of coded graphic characters is intended for use in data and text processing applications and also for information interchange. The set contains graphic characters used for general purpose applications in typical office environments in at least the following languages: English, Hebrew and Latin. It is not intended for pointed Hebrew. This set of coded graphic characters may be regarded as a version of an 8-bit code according to Standard ECMA-35 or Standard ECMA-43 at level 1. This ECMA Standard may not be used with any other ECMA Standards for 8-bit single-byte coded graphic character sets. If coded characters from more than one ECMA Standard are to be used together, by means of code extension techniques, the equivalent coded character sets from ISO/IEC 10367 should be used instead within a version of Standard ECMA-43 at level 2 or level 3. The coded characters in this set may be used in conjunction with coded control functions selected from ECMA-48. However, control functions are not used to create composite graphic symbols from two or more graphic characters (see clause 6). NOTE This ECMA Standard is not intended for use with Telematic services defined by ITU-T. If information coded according to this ECMA Standard is to be transferred to such services, it will have to conform to the requirements of those services at the access-point.
2
Conformance
2.1
Conformance of information interchange A coded-character-data-element (CC-data-element) within coded information for interchange is in conformance with this ECMA Standard if all the coded representations of graphic characters within that CC-data-element conform to the requirements of clause 6.
2.2
Conformance of devices A device is in conformance with this ECMA Standard if it conforms to the requirements of 2.2.1, and either or both of 2.2.2 and 2.2.3. A claim of conformance shall identify the document which contains the description specified in 2.2.1.
3
2.2.1
D e v ic e d e s c r ip t io n A device that conforms to this ECMA Standard shall be subject of a description that identifies the means by which the user may supply characters to the device, or may recognize them when they are made available to him, as specified respectively in 2.2.2 and 2.2.3.
2.2.2
Originating devices An originating device shall allow its user to supply any sequence of characters from those specified in clause 6, and shall be capable of transmitting their coded representations within a CC-data-element.
2.2.3
Receiving devices A receiving device shall be capable of receiving and interpreting any coded representations of characters that are within a CC-data-element, and that conform to clause 6, and shall make the corresponding characters available to its user in such a way that the user can identify them from among those specified there, and can distinguish them from each other.
References ECMA-6
7-Bit Input/Output Coded Character Set
ECMA-35
Code Extension Techniques
- 2 -
4
ECMA-43
8-Bit Coded Character Set Structure and Rules
ECMA-48
Control Functions for Coded Character Sets
ECMA-94
8-Bit Single-Byte Coded Graphic Character Sets - Latin Alphabets No. 1 to No. 4
ECMA-113
8-Bit Single-Byte Coded Graphic Character Sets - Latin/Cyrillic Alphabet
ECMA-114
8-Bit Single Byte Coded Graphic Character Sets - Latin/Arabic Alphabet
ECMA-118
8-Bit Single-Byte Coded Graphic Character Sets - Latin/Greek Alphabet
ECMA-128
8-Bit Single-Byte Coded Graphic Character Sets - Latin alphabet No. 5
ECMA-144
8-Bit Single-Byte Coded Graphic Character Sets - Latin Alphabet No. 6
Definitions For the purpose of this Standard the following definitions apply.
4.1
bi-directional text A text which may contain strings of characters with left-to-right and right-to-left directions.
4.2
bit combination An ordered set of bits used for the representation of characters.
4.3
byte A bit string that is operated upon as a unit.
4.4
character A member of a set of elements used for the organization, control, or representation of data.
4.5
code table A table showing the characters allocated to each bit combination in a code.
4.6
coded character set; code A set of unambiguous rules that establishes a character set and the one-to-one relationship between the characters of the set and their bit combinations.
4.7
coded-character-data-element (CC-data-element) An element of interchanged information that is specified to consist of a sequence of coded representations of characters, in accordance with one or more identified standards for coded character sets.
4.8
directional character properties A set of mutually exclusive properties which may qualify the members of a character set. These properties are used by algorithms which transform text from processing sequence into presentation sequence. Examples of values for directional character properties are "right-to-left", "left-to-right", "digit", "numeric separator", "neutral".
4.9
graphic character A character, other than a control function, that has a visual representation normally hand-written, printed or displayed, and that has a coded representation consisting of one or more bit combinations.
4.10
graphic symbol A visual representation of a graphic character or of a control function.
4.11
implicit directionality A text presentation method in which the direction is determined by an algorithm. The algorithm is based on the directional character properties of the character, its position relative to the preceding and following character and to the primary direction.
- 3 -
4.12
left-to-right character A character specific to a script written from left to right like the Latin script or the Greek script. Typical examples are the letters A to Z.
4.13
position That part of a code table identified by its column and row co-ordinates.
4.14
right-to-left character A character specific to a script written from right to left like the Arabic script or the Hebrew script. Typical examples are the letters of the Hebrew alphabet.
5 5.1
Notation, code table and names Notation The bits of the bit combinations of the 8-bit code are identified by b 8 , b 7 , b 6 , b 5 , b 4 , b 3 , b 2 and b 1 , where b 8 is the highest-order, or most-significant bit and b1 is the lowest-order, or least-significant bit. The bit combinations may be interpreted to represent numbers in binary notation by attributing the following weights to the individual bits: Bit
b8
b7
b6
b5
b4
b3
b2
b1
Weight
128
64
32
16
8
4
2
1
Using these weights, the bit combinations are identified by notations of the form xx/yy, where xx and yy are numbers in the range 00 to 15. The correspondence between the notations of the form xx/yy and the bit combinations consisting of the bits b 8 to b 1 is as follows: −
xx is the number represented by b8 , b 7 , b 6 and b 5 where these bits are given the weights 8, 4, 2, and 1, respectively.
−
yy is the number represented by b 4 , b 3 , b 2 and b 1 where these bits are given the weights 8, 4, 2, and 1, respectively.
The bit combinations are also identified by notations of the form hk, where h and k are numbers in the range 0 to F in hexadecimal notation. The number h is the same as the number xx described above, and the number k the same as the number yy described above.
5.2
Layout of the code table An 8-bit code table consists of 256 positions arranged in 16 columns and 16 rows. The columns and the rows are numbered 00 to 15. In hexadecimal notation the columns and the rows are numbered 0 to F. The code table positions are identified by notations of the form xx/yy, where xx is the column number and yy is the row number. The column and row numbers are shown at the top and left edges of the table, respectively. The code table positions are also identified by notations of the form hk, where h is the column number and k is the row number in hexadecimal notation. The column and row numbers are shown at the bottom and right edges of the table, respectively. The positions of the code table are in one-to-one correspondence with the bit combinations of the code. The notation of a code table position, of the form xx/yy, or of the form hk, is the same as that of the corresponding bit combination.
5.3
Names and meanings. This ECMA Standard assigns a unique name and a unique identifier to each graphic character. These names and identifiers have been taken from ISO/IEC 10646-1. This ECMA Standard also specifies an acronym for each of the characters SPACE, NO-BREAK SPACE, SOFT HYPHEN, LEFT-TO-RIGHT MARK and RIGHT-TO-LEFT-MARK. For acronyms only Latin capital letters A to Z are used. It is intended that the acronyms be retained in all translations of the text.
- 4 -
Except for SPACE (SP), NO-BREAK SPACE (NBSP), and SOFT HYPHEN (SHY), LEFT-TO-RIGHT MARK (LRM) and RIGHT-TO-LEFT MARK (RLM), this ECMA Standard does not define and does not restrict the meanings of graphic characters. This ECMA Standard specifies a graphic symbol for each graphic character. This symbol is shown in the corresponding position of the code table. However, this Standard does not specify a particular style or font design for imaging graphic characters. 5.3.1
SPACE (SP) A graphic character the visual representation of which consists of the absence of a graphic symbol.
5.3.2
NO-BREAK SPACE (NBSP) A graphic character the visual representation of which consists of the absence of a graphic symbol, for use when a line break is to be prevented in the text as presented.
5.3.3
S O F T H Y P H EN ( S H Y ) A graphic character that is imaged by a graphic symbol identical with, or similar to, that representing HYPHEN, for use when a line break has been established within a word.
5.3.4
LEFT-TO-RIGHT MARK (LRM) A graphic character the visual representation of which consists of the absence of a graphic symbol, which acts like a left-to-right character in a bi-directional text (such as LATIN SMALL LETTER A).
5.3.5
RIGHT-TO-LEFT MARK (RLM) A graphic character the visual representation of which consists of the absence of a graphic symbol, which acts like a right-to-left character in a bi-directional text (such as HEBREW LETTER ALEF).
6
Specification of the coded character set This ECMA Standard specifies 155 characters allocated to the bit combinations of the code table (table 2). Control functions, such as BACKSPACE or CARRIAGE RETURN, shall not be used to create composite graphic symbols, which are made up from the graphic representations of two or more characters.
6.1
Characters of the set and their coded representation See table 1.
- 5 -
Table 1 - Character set, coded representation Bit combination
Hex
Identifier
Name
02/00 02/01 02/02 02/03 02/04 02/05 02/06 02/07 02/08 02/09 02/10 02/11 02/12 02/13 02/14 02/15 03/00 03/01 03/02 03/03 03/04 03/05 03/06 03/07 03/08 03/09 03/10 03/11 03/12 03/13 03/14 03/15 04/00 04/01 04/02 04/03 04/04 04/05 04/06 04/07 04/08 04/09 04/10 04/11 04/12 04/13 04/14 04/15 05/00 05/01
20 21 22 23 24 25 26 27 28 29 2A 2B 2C 2D 2E 2F 30 31 32 33 34 35 36 37 38 39 3A 3B 3C 3D 3E 3F 40 41 42 43 44 45 46 47 48 49 4A 4B 4C 4D 4E 4F 50 51
U+0020 U+0021 U+0022 U+0023 U+0024 U+0025 U+0026 U+0027 U+0028 U+0029 U+002A U+002B U+002C U+002D U+002E U+002F U+0030 U+0031 U+0032 U+0033 U+0034 U+0035 U+0036 U+0037 U+0038 U+0039 U+003A U+003B U+003C U+003D U+003E U+003F U+0040 U+0041 U+0042 U+0043 U+0044 U+0045 U+0046 U+0047 U+0048 U+0049 U+004A U+004B U+004C U+004D U+004E U+004F U+0050 U+0051
SPACE EXCLAMATION MARK QUOTATION MARK NUMBER SIGN DOLLAR SIGN PERCENT SIGN AMPERSAND APOSTROPHE LEFT PARENTHESIS RIGHT PARENTHESIS ASTERISK PLUS SIGN COMMA HYPHEN-MINUS FULL STOP SOLIDUS DIGIT ZERO DIGIT ONE DIGIT TWO DIGIT THREE DIGIT FOUR DIGIT FIVE DIGIT SIX DIGIT SEVEN DIGIT EIGHT DIGIT NINE COLON SEMICOLON LESS-THAN SIGN EQUALS SIGN GREATER-THAN SIGN QUESTION MARK COMMERCIAL AT LATIN CAPITAL LETTER A LATIN CAPITAL LETTER B LATIN CAPITAL LETTER C LATIN CAPITAL LETTER D LATIN CAPITAL LETTER E LATIN CAPITAL LETTER F LATIN CAPITAL LETTER G LATIN CAPITAL LETTER H LATIN CAPITAL LETTER I LATIN CAPITAL LETTER J LATIN CAPITAL LETTER K LATIN CAPITAL LETTER L LATIN CAPITAL LETTER M LATIN CAPITAL LETTER N LATIN CAPITAL LETTER O LATIN CAPITAL LETTER P LATIN CAPITAL LETTER Q
- 6 -
Bit combination
Hex
Identifier
Name
05/02 05/03 05/04 05/05 05/06 05/07 05/08 05/09 05/10 05/11 05/12 05/13 05/14 05/15 06/00 06/01 06/02 06/03 06/04 06/05 06/06 06/07 06/08 06/09 06/10 06/11 06/12 06/13 06/14 06/15 07/00 07/01 07/02 07/03 07/04 07/05 07/06 07/07 07/08 07/09 07/10 07/11 07/12 07/13 07/14
52 53 54 55 56 57 58 59 5A 5B 5C 5D 5E 5F 60 61 62 63 64 65 66 67 68 69 6A 6B 6C 6D 6E 6F 70 71 72 73 74 75 76 77 78 79 7A 7B 7C 7D 7E
U+0052 U+0053 U+0054 U+0055 U+0056 U+0057 U+0058 U+0059 U+005A U+005B U+005C U+005D U+005E U+005F U+0060 U+0061 U+0062 U+0063 U+0064 U+0065 U+0066 U+0067 U+0068 U+0069 U+006A U+006B U+006C U+006D U+006E U+006F U+0070 U+0071 U+0072 U+0073 U+0074 U+0075 U+0076 U+0077 U+0078 U+007A U+007A U+007B U+007C U+007D U+007E
LATIN CAPITAL LETTER R LATIN CAPITAL LETTER S LATIN CAPITAL LETTER T LATIN CAPITAL LETTER U LATIN CAPITAL LETTER V LATIN CAPITAL LETTER W LATIN CAPITAL LETTER X LATIN CAPITAL LETTER Y LATIN CAPITAL LETTER Z LEFT SQUARE BRACKET REVERSE SOLIDUS RIGHT SQUARE BRACKET CIRCUMFLEX ACCENT LOW LINE GRAVE ACCENT LATIN SMALL LETTER A LATIN SMALL LETTER B LATIN SMALL LETTER C LATIN SMALL LETTER D LATIN SMALL LETTER E LATIN SMALL LETTER F LATIN SMALL LETTER G LATIN SMALL LETTER H LATIN SMALL LETTER I LATIN SMALL LETTER J LATIN SMALL LETTER K LATIN SMALL LETTER L LATIN SMALL LETTER M LATIN SMALL LETTER N LATIN SMALL LETTER O LATIN SMALL LETTER P LATIN SMALL LETTER Q LATIN SMALL LETTER R LATIN SMALL LETTER S LATIN SMALL LETTER T LATIN SMALL LETTER U LATIN SMALL LETTER V LATIN SMALL LETTER W LATIN SMALL LETTER X LATIN SMALL LETTER Y LATIN SMALL LETTER Z LEFT CURLY BRACKET VERTICAL LINE RIGHT CURLY BRACKET TILDE
10/00 10/01 10/02 10/03 10/04 10/05 10/06
A0 A1 A2 A3 A4 A5 A6
U+00A0
NO-BREAK SPACE (This position shall not be used) CENT SIGN POUND SIGN CURRENCY SIGN YEN SIGN BROKEN BAR
U+00A2 U+00A3 U+00A4 U+00A5 U+00A6
- 7 -
Bit combination
Hex
Identifier
Name
10/07 10/08 10/09 10/10 10/11 10/12 10/13 10/14 10/15 11/00 11/01 11/02 11/03 11/04 11/05 11/06 11/07 11/08 11/09 11/10 11/11 11/12 11/13 11/14 11/15 12/00 12/01 12/02 12/03 12/04 12/05 12/06 12/07 12/08 12/09 12/10 12/11 12/12 12/13 12/14 12/15 13/00 13/01 13/02 13/03 13/04 13/05 13/06 13/07 13/08 13/09 13/10 13/11
A7 A8 A9 AA AB AC AD AE AF B0 B1 B2 B3 B4 B5 B6 B7 B8 B9 BA BB BC BD BE BF C0 C1 C2 C3 C4 C5 C6 C7 C8 C9 CA CB CC CD CE CF D0 D1 D2 D3 D4 D5 D6 D7 D8 D9 DA DB
U+00A7 U+00A8 U+00A9 U+00D7 U+00AB U+00AC U+00AD U+00AE U+00AF U+00B0 U+00B1 U+00B2 U+00B3 U+00B4 U+00B5 U+00B6 U+00B7 U+00B8 U+00B9 U+00F7 U+00BB U+00BC U+00BD U+00BE
SECTION SIGN DIARESIS COPYRIGHT SIGN MULTIPLICATION SIGN LEFT-POINTING DOUBLE ANGLE QUOTATION MARK NOT SIGN SOFT HYPHEN REGISTERED SIGN MACRON DEGREE SIGN PLUS-MINUS SIGN SUPERSCRIPT TWO SUPERSCRIPT THREE ACUTE ACCENT MICRO SIGN PILCROW SIGN MIDDLE DOT CEDILLA SUPERSCRIPT ONE DIVISION SIGN RIGHT-POINTING DOUBLE ANGLE QUOTATION MARK VULGAR FRACTION ONE QUARTER VULGAR FRACTION ONE HALF VULGAR FRACTION THREE QUARTERS (This position shall not be used) (This position shall not be used) (This position shall not be used) (This position shall not be used) (This position shall not be used) (This position shall not be used) (This position shall not be used) (This position shall not be used) (This position shall not be used) (This position shall not be used) (This position shall not be used) (This position shall not be used) (This position shall not be used) (This position shall not be used) (This position shall not be used) (This position shall not be used) (This position shall not be used) (This position shall not be used) (This position shall not be used) (This position shall not be used) (This position shall not be used) (This position shall not be used) (This position shall not be used) (This position shall not be used) (This position shall not be used) (This position shall not be used) (This position shall not be used) (This position shall not be used) (This position shall not be used)
- 8 -
6.2
Bit combination
Hex
13/12 13/13 13/14 13/15 14/00 14/01 14/02 14/03 14/04 14/05 14/06 14/07 14/08 14/09 14/10 14/11 14/12 14/13 14/14 14/15 15/00 15/01 15/02 15/03 15/04 15/05 15/06 15/07 15/08 15/09 15/10 15/11 15/12 15/13 15/14 15/15
DC DD DE DF E0 E1 E2 E3 E4 E5 E6 E7 E8 E9 EA EB EC ED EE EF F0 F1 F2 F3 F4 F5 F6 F7 F8 F9 FA FB FB FD FE FF
Identifier
U+2017 U+05D0 U+05D1 U+05D2 U+05D3 U+05D4 U+05D5 U+05D6 U+05D7 U+05D7 U+05D9 U+05DA U+05DB U+05DC U+05DD U+05DE U+05DF U+05E0 U+05E1 U+05E2 U+05E3 U+05E4 U+05E5 U+05E6 U+05E7 U+05E8 U+05E9 U+05EA
U+200E U+200F
Name
(This position shall not be used) (This position shall not be used) (This position shall not be used) DOUBLE LOW LINE HEBREW LETTER ALEF HEBREW LETER BET HEBREW LETTER GIMEL HEBREW LETTER DALET HEBREW LETTER HE HEBREW LETTER VAV HEBREW LETTER ZAYIN HEBREW LETTER HET HEBREW LETTER TET HEBREW LETTER YOD HEBREW LETTER FINAL KAF HEBREW LETTER KAF HEBREW LETTER LAMED HEBREW LETTER FINAL MEM HEBREW LETTER MEM HEBREW LETTER FINAL NUN HEBREW LETTER NUN HEBREW LETTER SAMEKH HEBREW LETTER AYIN HEBREW LETTER FINAL PE HEBREW LETTER PE HEBREW LETTER FINAL TSADI HEBREW LETTER TSADI HEBREW LETTER QOF HEBREW LETTER RESH HEBREW LETTER SHIN HEBREW LETTER TAV (This position shall not be used) (This position shall not be used) LEFT-TO-RIGHT MARK RIGHT-TO-LEFT MARK (This position shall not be used)
Code table For each character in the set the code table (table 2) shows a graphic symbol at the position in the code table corresponding to the bit combination specified in table 1. The shaded positions in the code table correspond to bit combinations that do not represent graphic characters. Their use is outside the scope of this ECMA Standard; it is specified in other ECMA Standards, for example ECMA-48. The positions in the code table that are shown with cross-hatching correspond to bit combinations in table 1 having the entry "This position shall not be used".
- 9 -
Table 2 - Code table of Latin/Hebrew alphabet b8 0 0 0 0 0 0 0 0 1 1 1 1 1 1 1 1 b7 0 0 0 0 1 1 1 1 0 0 0 0 1 1 1 1 b6 0 0 1 1 0 0 1 1 0 0 1 1 0 0 1 1 b5 0 1 0 1 0 1 0 1 0 1 0 1 0 1 0 1
00 01 02 03 04 05 06 07 08 09 10 11 12 13 14 15 p SP NBSP 0 0 0 0 00 P 0 0
b4 b 3 b2 b1
0 0 0 1 01
1 A Q
a q
1
0 0 1 0 02
2 B R b r 3 C S c s
2
0 0 1 1 03
3 4
0 1 0 1 05
4 D T d t 5 E U e u
0 1 1 0 06
6 F V
6
1 0 0 0 08
f v 7 G W g w 8 H X h x
1 0 0 1 09
9
9 A
0 1 0 0 04
0 1 1 1 07
5 7 8
1 0 1 0 10
J Z
i y j z
1 0 1 1 11
K
k
B
1 1 0 0 12
L
C
1 1 0 1 13
M
l m
1 1 1 0 14
N
n
1 1 1 1 15
O _
o
SHY
LRM
D
RLM
E
x
1 2 3 4 5 6 7 8 9 A B C D E F
F he
0
I Y
99-0097-A
7 7.1
Identification of the character set Identification according to ECMA-35 and ECMA-43 The graphic characters of this ECMA Standard constitute a single coded character set. However, in accordance with ECMA-35 and ECMA-43 the code table of this ECMA Standard may be considered to consist of the following components: −
The character SPACE represented by bit combination 02/00;
−
a 94-character G0 graphic character set represented by bit combinations 02/01 to 07/14;
−
a 96-character G1 graphic character set represented by bit combinations 10/00 to 15/15.
When the identification methods of ECMA-35 or ECMA-43 are used, this ECMA Standard shall be identified by the following pair of designation functions:
- 10 -
GZD4
04/02
(ESC 02/08 04/02)
G1D6
04/07
(ESC 02/13 05/14)
NOTE The corresponding escape sequences are shown in parentheses.
7.2
Identification using the ISO International register of coded character sets to be used with escape sequences According to 7.1 above the character set of this ECMA Standard may be considered to consist of the character SPACE, a 94-character G0 graphic character set, and a 96-character G1 graphic character set. The G0 and G1 graphic character sets may be identified by the use of the Registration Numbers from the ISO International register of coded character sets to be used with escape sequences. When these registration numbers are used this ECMA Standard shall be identified by the following pair of registration numbers: −
G0 graphic character set ISO-IR 6
−
G1 graphic character set ISO/IR 198
- 11 -
Annex A ( in f o r ma tiv e )
Coverage of languages
A.1
Languages of European origin written in Latin script The following ECMA Standards specify coded character sets which comprise various different selections of characters based on the Latin alphabet. These sets are identified by the numbers 1 to 6 as shown: ECMA-94 ECMA-128 ECMA-144
Latin alphabets No. 1 to 4 Latin alphabet No. 5 Latin alphabet No. 6
The following official and regional languages written in Europe are covered by the Latin alphabets 1 to 6 as indicated by their number in table A.1:
Ta b le A . 1 - La n g u a g e c o v e r a g e Language
Covered by alphabet(s)
Albania Basque Breton Catalan Croat Czech Danish Dutch English Esperanto Estonian Faroese Finnish French
Language
1 2 5 Frisian 1 5 Galician 1 5 German 1 5 Greenlandic 2 Hungarian 2 Icelandic 1 4 5 6 Irish Gaelic (new orthography) 1 5 1 2 3 4 5 6 Italian 3 Latin 4 6 Latvian 1 6 Lithuanian 1 4 5 6 Luxemburgish (1) (3) (5) Maltese
Covered by alphabet(s) 1 5 1 5 1 2 3 4 5 1 4 5 2 1 1 5
Language
6 6 6 6
1 3 5 1 2 3 4 5 6 4 4 6 1 5 3
Norwegian Polish Portuguese Rhaeto-Romanic Romanian Sámi Scottish Gaelic Slovak Slovene Serbian Spanish Swedish Turkish
Covered by alphabet(s) 1
4 5 6 2
1 1
3
5 5
2 4 1 2 2 2 1 1
6 5
4
6
5 4 5 6 (3) 5
NOTES 1.
The list of languages in table A.1 is not exhaustive. It shows the languages that are included in the Scope clause of the Latin alphabets.
2.
For writing French, three characters (Œ, œ, Ÿ) not specified in Latin alphabets 1, 3 and 5, are also needed.
3.
The various Sámi languages use partly differing orthographies. The character sets in Latin alphabets No. 4 and No. 6 cover the requirements of the Sámi languages most commonly used in Finland, Norway and Sweden. For the Skolt Sámi language used in Finland and Norway additional characters are needed.
4.
There are several official written languages outside Europe that are covered by Latin alphabet No. 1. Examples are Indonesian/Malay, Tagalog (Philippines), Swahili, Afrikaans.
5.
Use of Latin alphabet No. 3 for Turkish is deprecated.
- 12 -
A.2
Languages written in non-Latin scripts The following standards specify coded character sets which include graphic characters from alphabets other than the Latin alphabet: ECMA-113 ECMA-114 ECMA-118 ECMA-121
Latin/Cyrillic alphabet Latin/Arabic alphabet Latin/Greek alphabet Latin/Hebrew alphabet
The following official and regional languages are covered by these alphabets: Cyrillic characters included in Standard ECMA-113 cover Bulgarian, Byelorussian, (Slavic) Macedonian, Russian, Serbian and Ukranian (as written up to 1990, see also the Scope of Standard ECMA-113). The Arabic characters included in .Standard ECMA-114 cover Arabic. The Greek characters included in ECMA-118 cover Greek (monotonikó orthography). The Hebrew characters included in ECMA-121 cover Hebrew.
- 13 -
Annex B ( in f o r ma tiv e )
Main differences between the first edition and this second edition of ECMA-121
B.1
The names of the graphic characters have been amended where necessary to align them with the names of the characters adopted for all standards on coded character sets developed under the responsibility of ISO/IEC JTC 1. For each character the short identifiers specified in ISO/IEC 10646-1, Amendment 9, have been added to table 1.
B.2
The new style of conformance clause, adopted for all standards on coded character sets, has been introduced.
B.3
Object identifiers conforming to Abstract Syntax Notation One (ASN.1, see ISO/IEC 8824-1) are specified in annex E for the character set, and the corresponding coded representations of this ECMA Standard. Registration numbers from the International register of coded character sets to be used with escape sequences have been included as an additional method of identifying the coded character set of this ECMA Standard.
B.4
A new annex A has been added that identifies the coverage of languages by all Latin alphabets.
B.5
Various editorial adjustments and clarifications have been made to the text of the Standard. The hexadecimal equivalents of the bit combinations have been added to tables 1 and 2.
B.6
Support for bi-directionality has been included and is described in annex C.
B.7
Annex D, Bibliography, has been added.
- 14 -
- 15 -
Annex C ( in f o r ma tiv e )
Bi-directional text support
C.1
Bi-directional text formatting The LEFT-TO-RIGHT MARK and RIGHT-TO-LEFT MARK characters are used in formatting bi-directional text. Text conforming to this ECMA Standard is typically rendered with an implicit bi-directional algorithm. An implicit algorithm uses the directional character properties to determine the correct display order of characters on a horizontal line of text. The characters specified in 5.3.4 and 5.3.5 are acting exactly like left-to-right or right-to-left characters in terms of affecting ordering (bi-directional format marks). They have no visible graphic symbols, and they do not have any other semantic effect. An algorithm supporting bi-directional text formatting is described in the Unicode standard. This ECMA Standard does not preclude the use of other means to manage the text format, such as the use of control functions from ECMA-48, or other means external to the graphic character set.
C.2
Backward compatibility The two characters LEFT-TO-RIGHT MARK and RIGHT-TO-LEFT MARK are used in applications that have a bi-directional capability, i.e. applications supporting implicit directionality. The behaviour of applications not supporting implicit directionality upon receiving of these characters is outside the scope of this ECMA Standard and therefore undefined.
- 16 -
- 17 -
Annex D ( in f o r ma tiv e )
Bibliography
ECMA-48
Control Functions for Coded Character Sets (1991)
ECMA TR/53
Handling of bi-directional texts (1992)
ISO/IEC 10646-1:1993 - Information technology -Universal Multiple-Octet Coded Character Set (UCS) - Part 1: Architecture and Basic Multilingual Plane ISO International register of coded character sets to be used with escape sequences The Standards Institution of Israel: SI 1311 (July 1989 - in revision), Information technology - ISO 8-bit coded character set for information interchange The Unicode Consortium, The Unicode Standard - Version 2.0 (1996).
Free printed copies can be ordered from: ECMA 114 Rue du Rhône CH-1204 Geneva Switzerland Fax: Email:
+41 22 849.60.01 [email protected]
Files of this Standard can be freely downloaded from the ECMA web site (www.ecma.ch). This site gives full information on ECMA, ECMA activities, ECMA Standards and Technical Reports.
ECMA 114 Rue du Rhône CH-1204 Geneva Switzerland See inside cover page for obtaining further soft or hard copies.