Code page 949 (IBM)

From Wikipedia, the free encyclopedia
  (Redirected from Code page 944)
Jump to navigation Jump to search
IBM code page 949
IBM Korea KS PC-Data encoding (IBM-949).svg
Layout of the IBM-949 code page
Alias(es)
  • IBM-949, x-IBM949
  • ASCII-based: IBM-949C, x-IBM949C, cp949c
  • Ambiguous with UHC: 949, cp949
Language(s)Korean
Created byIBM
ClassificationExtended ISO 646, variable-width encoding, CJK encoding
ExtendsEUC-KR
Preceded byCode page 944

IBM code page 949 (IBM-949) is a character encoding which has been used by IBM to represent Korean language text on computers. It is a variable-width encoding which represents the characters from the Wansung code defined by the South Korean standard KS X 1001 in a format compatible with EUC-KR, but adds IBM extensions for additional hanja, additional precomposed Hangul syllables, and user-defined characters.

Giving values in hexadecimal, bytes 0x00 through 0x7F are used for single byte KS X 1003 (ISO 646:KR) characters, a similar set to ASCII but with a won sign rather than a backslash. Bytes 0x80 through 0x84 are used for IBM single byte extension characters. Lead bytes 0x8F through 0xA0 are used for IBM double byte extension characters. Lead bytes 0xA1 through 0xFE are used for Wansung code (KS X 1001 characters in EUC-KR form, double byte), but with some unused space opened up for user-defined use.

Although both are sometimes named "cp949", IBM-949 is different from Windows code page 949 (IBM-1363), which is Microsoft's Unified Hangul Code, a different extension of EUC-KR. It should also not be confused with IBM's implementation of plain EUC-KR (IBM-970).

Terminology and encoding labelling[edit]

Both IBM-949 and Unified Hangul Code (Windows-949) are known as "code page 949" (or "cp949") although they share only the EUC-KR subset in common. Neither has a standardised IANA-registered label to identify it. Although UHC is included in the WHATWG Encoding Standard,[1] with labels including "windows-949",[2] IBM-949 is not. IBM-949 therefore is not permitted in HTML5.

Although the meaning of the label "ibm-949" (and conversely "windows-949" and "ms949") is unambiguous where these labels are supported, the interpretation of the encoding labels "949" and "cp949" consequently varies between implementations. For example, International Components for Unicode uses "cp949", "949", "ibm-949" and "x-IBM949" to refer to IBM-949,[3] and additionally the labels "cp949c", "ibm-949c" and "x-IBM949C" to refer to an variant which uses unmodified ASCII mappings for 0x20–7E (resulting in duplicate mappings for the backslash),[4] while (of the labels incorporating the code page number 949) only "ms949" and "windows-949" are assigned to UHC.[5] This is in contrast to Python, which recognises both "cp949" and "949" (in addition to the more explicit "ms949" and "uhc", but not "windows-949") as labels for UHC, and does not include an IBM-949 codec.[6]

IBM-949 is a variable width encoding defined as the combination of two fixed-width code pages, the single-byte Code page 1088 and the double-byte Code page 951.[7][8][9]

History[edit]

A version of Code page 951 (a DBCS-PC, i.e. double-byte non-EUC non-EBCDIC, code), the double-byte component for IBM-949, is defined in the September 1992 revision of IBM Corporate Specification C-H 3-3220-125, along with Code page 834 (a DBCS-Host, i.e. double-byte EBCDIC, code), which is the double byte component of Code page 933.[10] This version of Code page 949/951 considered the entire lead byte range 0x8F–A0 to be a user-defined region, and included only standard Wansung assignments and user-defined areas, thus not including some characters which Code page 933/834 included.[10] Some later versions, such as that implemented by International Components for Unicode (ICU), shrink the user-defined region to include these characters as extensions.[11]

IBM code pages 934 and 944
Language(s)Korean
ExtendsN-byte Hangul Code
Transforms / EncodesCode page 933
Succeeded byIBM code page 949

The earlier October 1989 revision of C-H 3-3220-125 had instead defined Code page 926 as its DBCS-PC code, which encoded the same characters as IBM-834 in a layout differing from both IBM-951 and IBM-834, which had a different lead byte range and was not an EUC-KR extension.[10] IBM-926 was combined with Code page 891 or Code page 1040 (respectively 8-bit N-byte Hangul Code and an extension thereof; compare how Shift JIS extends 8-bit JIS X 0201) to form IBM-934 or IBM-944 respectively.[12][13]

Code page 944/926 are now deprecated in favour of IBM-949. The 1992 revision designates code page 926 as "restricted" ("limited to the particular environment for which [it is] registered") and does not give its chart or mappings from the other code pages,[10] and CCSID 944 is categorised as "coexistence and migration"[13] (contrast "interoperable" for CCSID 949).[7] International Components for Unicode includes Unicode mappings for IBM-949[3][11] and IBM-933, but its IBM-944 mapping was removed in 2001.[14]

Single byte codes[edit]

IBM code page 949 (single byte component: 1088)[15][16][3][4][11]
_0 _1 _2 _3 _4 _5 _6 _7 _8 _9 _A _B _C _D _E _F
0_
0
NUL
0000

250C

2510

2514

2518

2502

2500

2022

25D8

25CB

25D9

2642

2640

266A

266B

263C
1_
16

253C

25C4

2195

203C

00B6

2534

252C

2524

2191

251C

2192

2190

221F

2194

25B2

25BC
2_
32
SP
0020
!
0021
"
0022
#
0023
$
0024
%
0025
&
0026
'
0027
(
0028
)
0029
*
002A
+
002B
,
002C
-
002D
.
002E
/
002F
3_
48
0
0030
1
0031
2
0032
3
0033
4
0034
5
0035
6
0036
7
0037
8
0038
9
0039
:
003A
;
003B
<
003C
=
003D
>
003E
?
003F
4_
64
@
0040
A
0041
B
0042
C
0043
D
0044
E
0045
F
0046
G
0047
H
0048
I
0049
J
004A
K
004B
L
004C
M
004D
N
004E
O
004F
5_
80
P
0050
Q
0051
R
0052
S
0053
T
0054
U
0055
V
0056
W
0057
X
0058
Y
0059
Z
005A
[
005B

20A9
]
005D
^
005E
_
005F
6_
96
`
0060
a
0061
b
0062
c
0063
d
0064
e
0065
f
0066
g
0067
h
0068
i
0069
j
006A
k
006B
l
006C
m
006D
n
006E
o
006F
7_
112
p
0070
q
0071
r
0072
s
0073
t
0074
u
0075
v
0076
w
0077
x
0078
y
0079
z
007A
{
007B
|
007C
}
007D
~
007E

2302
8_
128
¢
00A2
¬
00AC
\
005C

203E
¦
00A6
UDC
LEAD
9_ UDC
LEAD
UDC
LEAD
UDC
LEAD
UDC
LEAD
UDC
LEAD
UDC
LEAD
UDC
LEAD
UDC
LEAD
UDC
LEAD
UDC
LEAD
Hanja/Sym
LEAD
Ext Hanja
LEAD
Ext Hanja
LEAD
Ext Hanja
LEAD
Hanja/Syll
LEAD
Ext Syll
LEAD
A_ Ext Syll
LEAD
 
Punct
LEAD
1-_
Symbol
LEAD
2-_
ISO646
LEAD
3-_
Jamo
LEAD
4-_
Greek
LEAD
5-_
Box
LEAD
6-_
Units
LEAD
7-_
Ext Alnum
LEAD
8-_
Ext Alnum
LEAD
9-_
Hiragana
LEAD
10-_
Katakana
LEAD
11-_
Cyrillic
LEAD
12-_
 
 
13-_
 
 
14-_
 
 
15-_
B_ Syllable
LEAD
16-_
Syllable
LEAD
17-_
Syllable
LEAD
18-_
Syllable
LEAD
19-_
Syllable
LEAD
20-_
Syllable
LEAD
21-_
Syllable
LEAD
22-_
Syllable
LEAD
23-_
Syllable
LEAD
24-_
Syllable
LEAD
25-_
Syllable
LEAD
26-_
Syllable
LEAD
27-_
Syllable
LEAD
28-_
Syllable
LEAD
29-_
Syllable
LEAD
30-_
Syllable
LEAD
31-_
C_ Syllable
LEAD
32-_
Syllable
LEAD
33-_
Syllable
LEAD
34-_
Syllable
LEAD
35-_
Syllable
LEAD
36-_
Syllable
LEAD
37-_
Syllable
LEAD
38-_
Syllable
LEAD
39-_
Syllable
LEAD
40-_
UDC
LEAD
41-_
Hanja
LEAD
42-_
Hanja
LEAD
43-_
Hanja
LEAD
44-_
Hanja
LEAD
45-_
Hanja
LEAD
46-_
Hanja
LEAD
47-_
D_ Hanja
LEAD
48-_
Hanja
LEAD
49-_
Hanja
LEAD
50-_
Hanja
LEAD
51-_
Hanja
LEAD
52-_
Hanja
LEAD
53-_
Hanja
LEAD
54-_
Hanja
LEAD
55-_
Hanja
LEAD
56-_
Hanja
LEAD
57-_
Hanja
LEAD
58-_
Hanja
LEAD
59-_
Hanja
LEAD
60-_
Hanja
LEAD
61-_
Hanja
LEAD
62-_
Hanja
LEAD
63-_
E_ Hanja
LEAD
64-_
Hanja
LEAD
65-_
Hanja
LEAD
66-_
Hanja
LEAD
67-_
Hanja
LEAD
68-_
Hanja
LEAD
69-_
Hanja
LEAD
70-_
Hanja
LEAD
71-_
Hanja
LEAD
72-_
Hanja
LEAD
73-_
Hanja
LEAD
74-_
Hanja
LEAD
75-_
Hanja
LEAD
76-_
Hanja
LEAD
77-_
Hanja
LEAD
78-_
Hanja
LEAD
79-_
F_ Hanja
LEAD
80-_
Hanja
LEAD
81-_
Hanja
LEAD
82-_
Hanja
LEAD
83-_
Hanja
LEAD
84-_
Hanja
LEAD
85-_
Hanja
LEAD
86-_
Hanja
LEAD
87-_
Hanja
LEAD
88-_
Hanja
LEAD
89-_
Hanja
LEAD
90-_
Hanja
LEAD
91-_
Hanja
LEAD
92-_
Hanja
LEAD
93-_
UDC
LEAD
94-_
 
 
 
_0 _1 _2 _3 _4 _5 _6 _7 _8 _9 _A _B _C _D _E _F

  Letter  Number  Punctuation  Symbol  Other  Lead byte  Undefined  Differences from code page 437 (for 0x00–7F) or EUC-KR (for 0x80–FF)

Double byte codes[edit]

Lead bytes 0x8F–99, 0xC9, 0xFE (user defined ranges)[edit]

IBM-949 is designed to support a maximum of 1880 UDC (user-defined characters),[7] including ranges within unused rows of the Wansung plane, and ranges outside the Wansung plane. In this version, the lead bytes 0x8F–A0 contain a maximum of 1692 UDC, and lead bytes 0xC9 and 0xFE contain a maximum of 94 each (i.e. with trail bytes 0xA0–FE).[10] However, when the extensions to support the entire double-byte repertoire of IBM-933 are implemented, they use lead bytes 0x9A–A0, resulting in a smaller maximum number of characters left for user definition.[3][11]

When mapped to Unicode, 0xC9A1–C9FE (between the syllable and hanja ranges) are mapped to the Unicode Private Use Area code points U+E000–E05D, while 0xFEA1–FEFE (between the end of the hanja range and the end of the plane) are mapped to U+E05E–E0BB. Outside the Wansung plane, 0x8FA0–9AA5 (where the second byte is in the range 0xA1–FE) are mapped to the Private Use Area code points U+E0BC–E4CA.[3] The last of these ranges cuts into the start of the 0x9A row (shown below).

Collectively these private use ranges cover the code points U+E000..E4CA, allowing 1227 UDC to be mapped from IBM-949 to Unicode.[11] The separate private use area range U+F843..F86E is used by IBM to map some characters within the extended hanja range.[11] This follows early recommendations from the Unicode Consortium that corporate characters be allocated from U+F8FF downward and user-defined characters be allocated from U+E000 upward,[17] and is part of a larger corporate private use area scheme which is defined internally by IBM, and includes 192 characters and three unused positions in the range U+F83D..F8FF.[18]

Lead bytes 0x9A–9D (extended symbols and hanja)[edit]

0x9AA1 through 0x9AA5 are the end of the user-defined range. The remainder of this range includes some non-Hangul characters included in Code page 933 but not in Wansung code. 0x9AA6 through 0x9AAB contain miscellaneous technical or mathematical symbols. The remainder contains hanja additional to those included in KS X 1001, although some are mapped by IBM to the Private Use Area.

IBM code page 949 (prefixed with 0x9A)[11][20]
_0 _1 _2 _3 _4 _5 _6 _7 _8 _9 _A _B _C _D _E _F
A_ PUA
E4C6
PUA
E4C7
PUA
E4C8
PUA
E4C9
PUA
E4CA
ǂ[a]
01C2

2266

2267

212A

FFE4
ʺ
02BA

F843/5580

64F1

7FAF

9163
B_
F844/91B5

9ABC

84B9

54FD

6243

6AA0

71B2

754A

F845/7A27

96DE

6772

77BD

8A41

6831

69D3

7B9C
C_
874C

970D

76E5

9E1B

9278

4F5D

50B4

5ABE

5AD7

F846/6677

750C

89AF

98B6

63AC

8DEA

5DF9
D_
6F0C

5C8C

7B08

F847/8987

9C2D

F848/551C

7CEF

5583

66E9

8FFA

4F5E

F849/7370

5B65

F84A/9B27

977C

601B
E_
95E5

97C3

515A

87F7

7893

83DF

5484

578C

809A

86AA

6ED5

706F

9419

7296

5E71

57D3
F_
6994

6DBC

9B4E

7658

8182

8821

9462

6ADF

9B23

6624

6CE0

82D3

86C9

6F66

826B
IBM code page 949 (prefixed with 0x9B)[11][20]
_0 _1 _2 _3 _4 _5 _6 _7 _8 _9 _A _B _C _D _E _F
A_
8F64

6F09
祿
F84B/797F

8F46

7C5F

857E

8A84

F84C/5BE5

50C2

9ACF
窿
7ABF

51DB

5EE9

F84D/63D0

6F13
B_
79BB

87AD

9B51

75F3

5CA6

5ABD

87C7

8B3E

93DD

9B18

9B4D

771B

82FA

8109

4FDB

8004
C_
927E

6FDB

77C7

7030

7CDC

95A9

F84E/5A46

6B02

7254

80D6

9AE3

9B74

F84F/6F58

7FFB

8F9F

6C74
D_
8FAE

F850/904D

99E2

5F46

8FF8

9D07

9EFC

8760

4E30

8451

4EC6

7F58

82FB

8709

982B

9B92
E_
F851/541F

8561
巿
F852/5DFF

9AF4

9EFB

59A3

F853/6C99

6C98

7765

7BE6

8153

8F61

9AC0

64EF

860B

F854/8D07
F_
9870

9B22

59D2

F855/9E9D

6942

69CE

7B25

69CA

9460

6B43

9364

970E

6BA4

9C13

566C
IBM code page 949 (prefixed with 0x9C)[11][20]
_0 _1 _2 _3 _4 _5 _6 _7 _8 _9 _A _B _C _D _E _F
A_ 婿
5A7F

F856/9F5F

F857/5C04

F858/55AE

5C20

6103

F859/6D17

71F9

F85A/9730

5070

F85B/F909

6308

8258

9704

87C0
B_
7463

53DF
宿
F85C/5BBF

666C

6EB2

795F

F85D/96CE

9D89

8671

557B

F85E/5BFA

7DE6

77E7

F85F/745F

843C

8D0B
C_
9D08

621E

904F

5D52

8AF3

9EEF

9785

6B38

769A

7919

9749

9628

F860/5C04

7BDB

7C65

7F98
D_
6554

605A

F861/5C04

F862/7FA8

81D9

8815

8B8C

5869

995C

5B30

7768

7FF3

854B

9068

5ABC

8580
E_
9C2E

8558

8202

86F9

5401

71A8

873F

5E43

885E

56FF

5E37

8564

9EDD

9B3B

6ABC

73E2
F_
9F66

6339

682E

9823

4EDE

7725

7CA2

8014

89DC

8D6D

67DE

6F5C

8695

5D82

7634
IBM code page 949 (prefixed with 0x9D)[11][20]
_0 _1 _2 _3 _4 _5 _6 _7 _8 _9 _A _B _C _D _E _F
A_
88C5

7E94

67E2

86C6

8C6C

7CF4

56C0

5DD3

78DA

7FE6

7A83

6904

6883

6662

7445
B_
8E36

F863/540A

566A

7681

7AC8

7B0A

7CF6

7D5B

9BDB

6A05

8E64

851F

8098

96BC

F864/5247

8A3C
C_
75E3

F865/6E4C

615A

5231

60B5

6C05

7C00

8734

8E91

6FFA

7C37

873B

780C

9746

5CED

7D83
D_
9214

9798

F866/6578

8E85

9AD1

6031

8471

6467

F867/69CC

7503

7B92

97A6

9E81

9EA4

F868/677B

8233
E_
51B2

6A47

F869/8D05

5DF5

5FB4

9D44

5FF1

62C6

6A50

99C4

F86A/5E40

8759

5E96

70AE

8216

924B
F_
9784

F86B/5206

84D6

8E55

7627

F86C/90AF

9DF3

7095

5EE8

F86D/614A

7BCB

F86E/965C

769E

9190

9DBB

Lead bytes 0x9E–A0 (extended hanja and syllables)[edit]

0x9EA1 through 0x9EAC contain the remainder of the extended hanja. The rest of the range contains a few additional Hangul syllables which are not available in pre-composed form in pure EUC-KR. Unlike Unified Hangul Code, this is insufficient to support all non-partial Johab syllables absent in Wansung code.

Significant amongst these are 뢔 (rwae, 0x9EFC), 쌰 (ssya, 0x9FE6), 쎼 (ssye, 0x9FED), 쓔 (ssyu, 0x9FF3) and 쬬 (jjyo, 0xA0C1), which correspond to the beginnings of the standard Wansung characters 뢨, 썅, 쏀, 쓩, and 쭁 respectively, when partly entered in an input method editor.

IBM code page 949 (prefixed with 0x9E)[11]
_0 _1 _2 _3 _4 _5 _6 _7 _8 _9 _A _B _C _D _E _F
A_
944A

571C

61FD

9B1F

5A93

6033

56C2

7334

7BCC

5FFB

8FC4

9821

AC02

AC0B

AC79
B_
AC87

AC93

ACE9

ACFA

AD19

AD28

AD2B

AD9B

ADD5

ADEC

AE02

AE0F

AE11

AE27

AE3C

AE44
C_
AE49

AE62

AEA0

AF04

AF33

AF4C

AF58

AF5B

AF68

AF93

AFB2
꾿
AFBF

AFD8

AFE7

B00D

B021
D_
B060

B090

B0BB

B0EC

B10F

B11E

B147

B153

B159

B16F

B17A

B1A7

B1B0

B233

B2A7

B2C1
E_
B2D1

B2E0

B331

B338

B368

B36A

B39C

B3D3

B400

B40F

B42C

B457

B47F

B4B4

B4C1

B4E7
F_
B52E

B532

B537

B53F

B568

B584

B5F4

B680

B6B8

B70C

B7D0

B80F

B894

B8DC

B917
IBM code page 949 (prefixed with 0x9F)[11]
_0 _1 _2 _3 _4 _5 _6 _7 _8 _9 _A _B _C _D _E _F
A_
B990

B9DF

B9FB

BA1C

BA6B

BA6D

BA80

BAAF

BAC3

BAE0

BBC1

BBD5

BBDC

BBE0

BC0E
B_
BC28

BC37

BC5C

BC68

BC98

BC9C

BCB9

BCCC

BCD2

BCD3

BCD4

BD23

BD97

BDB4

BE18

BE21
C_
BE28

BE4B

BE9C

BEB4

BEED

BEF0

BEF4
뻿
BEFF

BF24

BF5C

BF78

BFC0

BFD5

BFDD

BFE8

C004
D_
C020

C059

C074

C0AE

C0B7

C0BB

C0C3

C0C7

C0CF

C125

C13F

C151

C157

C193

C1BB

C28C
E_
C2B3

C2C0

C2E6

C302

C30B

C327

C330

C343

C34C

C37B

C385

C399

C3A0

C3BC

C3FC

C43F
F_
C477

C493

C4D3

C4D4

C53C

C53F

C54F

C55F

C590

C5AB

C5B6

C5F1

C5F3

C61D

C62B
IBM code page 949 (prefixed with 0xA0)[11]
_0 _1 _2 _3 _4 _5 _6 _7 _8 _9 _A _B _C _D _E _F
A_
C63A

C6B7

C6DF

C70B

C736

C77B

C7A7

C7AA

C807

C814

C81B

C839

C84B

C890

C89C
B_
C8A0

C8AC

C8B0

C8B8

C8E8

C8F0

C8F1

C92B

C96D

C9A4

C9D4

CA30

CA57

CA70

CA97

CAA0
C_
CAD2

CB2C

CB80

CBE5

CBF0

CC1F

CC26

CC2F
찿
CC3F

CC42

CC71

CC7C

CCE3

CCE5

CD40

CDC3
D_
CE3C

CE7B

CE97

CEA9

CEC8

CEFD

CF19

CF8D

CF90

CF9F

CFAC

CFBD

D088

D114

D160

D169
E_
D277

D293

D2CD

D2E7

D30A

D326

D359

D360

D3B2

D3B5

D3C7

D424

D4B0

D4E9

D505

D510
F_
D519

D520

D524

D55F

D561

D56C

D571

D5AC

D5CF

D62C

D6A9

D6B8

D6D5

D70C

D76D

Lead bytes 0xA1–C8, 0xCA–FD (standard Wansung)[edit]

See also[edit]

Footnotes[edit]

  1. ^ This is not included for IPA support. Rather, in Code page 933, SO 0x4160 is a not-equals sign displayed with a slash, while IBM-933 SO 0x418D is one displayed with a backslash (i.e. =⃥).[10] Although it is IBM-933 SO 0x4160 which is mapped to the usual not-equals GCGID SA540080 (fullwidth of SA540000), it is IBM-933 SO 0x418D which is mapped to EUC-KR and IBM-949 0xA1C1,[10] due to the reference glyph for the not-equals sign in KS C 5601-1987 also showing it with a backslash.[21] Hence, U+2260, which is mapped to EUC-KR and therefore IBM-949 0xA1C1, is mapped to IBM-933 SO 0x418D, leaving IBM-933 SO 0x4160 (and therefore IBM-949 0x9AA6) to be mapped to the visually similar character at U+01C2.[22]

References[edit]

  1. ^ van Kesteren, Anne, "5. Indexes (§ index EUC-KR)", Encoding Standard, WHATWG, This matches the KS X 1001 standard and the Unified Hangul Code, more commonly known together as Windows Codepage 949.
  2. ^ van Kesteren, Anne. "4.2. Names and labels". Encoding Standard. WHATWG.
  3. ^ a b c d e "Converter Explorer: ibm-949_P110-1999 (alias x-IBM949)", International Components for Unicode, Unicode Consortium
  4. ^ a b "Converter Explorer: ibm-949_P11A-1999 (alias x-IBM949C)", International Components for Unicode, Unicode Consortium. This is the ASCII-based version of IBM-949.
  5. ^ "windows-949-2000", Converter Explorer, International Components for Unicode
  6. ^ "codecs — Codec registry and base classes § Standard Encodings". Python 3.7.2 documentation. Python Software Foundation.
  7. ^ a b c "Coded character set identifiers: CCSID 949". IBM Globalization. IBM. Archived from the original on 2014-11-29.
  8. ^ "CCSID 1088 information document". Archived from the original on 2016-03-26.
  9. ^ "Code page 951 information document". Archived from the original on 2017-01-16.
  10. ^ a b c d e f g h "IBM Korean Graphic Character Set: DBCS-Host and DBCS-PC" (PDF). IBM. 2001 [1992]. C-H 3-3220-125 1992-09.
  11. ^ a b c d e f g h i j k l m International Components for Unicode (ICU), ibm-949_P110-1999.ucm, 2002-12-03
  12. ^ "Coded character set identifiers: CCSID 934". IBM Globalization. IBM. Archived from the original on 2014-12-02.
  13. ^ a b "Coded character set identifiers: CCSID 944". IBM Globalization. IBM. Archived from the original on 2014-12-01.
  14. ^ Viswanadha, Ram (2001-11-01). "ICU-1281 Remove unwanted ucmfiles". International Components for Unicode.
  15. ^ Code Page CPGID 01088 (pdf) (PDF), IBM
  16. ^ Code Page CPGID 01088 (txt), IBM
  17. ^ "2.0: Changes in Unicode 1.0" (PDF). The Unicode Standard, Version 1.1. Unicode Consortium. pp. 3–4. UTR #4.
  18. ^ a b "CPGID 01449: IBM default PUA". IBM Globalization: Code page identifiers. IBM. Archived from the original on 2015-09-16. IBM has designated 195 positions from U+F83D to U+F8FF for use as IBM Corporate-zone and intends to use them consistently within IBM whenever there is a need to maintain the round-trip integrity of IBM characters. […] At present CS 3099 containing 192 IBM Corportate [sic] characters has been defined.
  19. ^ "ibm-933_P110-1995.ucm". International Components for Unicode.
  20. ^ a b c d Private Use Area mapped hanja are identified from code charts. The IBM document C-H 3-3220-125 1992-09 gives code charts for the code pages used as the double-byte components for Code page 933 and an older version of Code page 949 without these extensions; however, the hanja in this section correspond to (and are in the same order as) the subset of table 7 for which a "PC Code" is not listed.[10] The Corporate Private Use Area mappings are also co-ordinated with other code pages,[18] including Code page 933,[19] which can be used to obtain the "Host Code" for a given Corporate Private Use Area mapping.
  21. ^ Korea Bureau of Standards (1988-10-01). Korean Graphic Character Set for Information Interchange (PDF). ITSCJ/IPSJ. ISO-IR-149.
  22. ^ "ibm-933_P110-1995 (lead bytes 0E41)". Converter Explorer. International Components for Unicode.