女人的逼什么样| 3.30是什么星座| 金刚是什么意思| 一个九一个鸟念什么| 宫颈液基细胞学检查是什么| 大便有点绿色是什么原因| 血沉低是什么意思| 乳糖不耐受可以喝什么奶| 且慢是什么意思| 李子什么时候吃最好| 肝胆挂什么科| 肌酐高不能吃什么| 霉菌有什么症状| 焗油是什么意思| 手汗症挂什么科| 小孩经常尿床是什么原因| 什么是青光眼| 17度穿什么衣服合适| 水漂是什么意思| 陕西什么面| 为什么老是犯困想睡觉| 3的倒数是什么| 房产证改名字需要什么手续| 什么才是真正的情人| 发烧为什么不能吃鸡蛋| 乳腺结节低回声是什么意思| 什么叫夏至| 指滑是什么意思| 加湿器有什么用| 助听器什么牌子的好| 二椅子什么意思| 为什么屁股上会长痘| 子欲养而亲不待是什么意思| 面子是什么意思| 男人不够硬吃什么好| 不成功便成仁的仁是什么意思| 尿淀粉酶高是什么原因| 姓陆的女孩取什么名字好| 馥字五行属什么| 周易和易经有什么区别| 飞蓬草有什么功效| 什么是平舌音| 什么是情人| 颢字五行属什么| 尿酸高要吃什么药| 翼龙吃什么| 98年虎是什么命| 风象星座是什么意思| 舌头白腻厚苔是什么原因| 低压高吃什么药效果好| 消化不良用什么药| 天河水是什么意思| 澳门使用什么货币| 1月22号什么星座| 雷锋属什么生肖| 白术有什么作用| 中元节不能穿什么衣服| 安然无恙是什么意思| 为什么会感染幽门螺杆菌| 周中是什么意思| 肝功高是什么原因引起的| 男朋友生日送什么礼物| 狐假虎威是什么意思| 杆菌是什么意思| 老夫老妻什么意思| 低频是什么意思| 坛城是什么意思| 安静如鸡什么意思| 戴字五行属什么| ca125检查是什么意思| 押韵什么意思| 有一种水果叫什么竹| 烧心是什么原因引起的| 八纲辨证中的八纲是什么| 自然堂适合什么年龄| 王朝马汉是什么意思| 什么样的人容易中风| 什么是白噪音| i是什么| 姐姐的孩子叫什么| 孩子肚子疼吃什么药| 饮水思源是什么意思| 为什么会感染幽门螺杆菌| 水逆退散是什么意思| 才高八斗什么生肖| 头皮屑是什么| 直系亲属为什么不能输血| 岳飞属什么生肖| 皮下出血点是什么原因| 开普拉多的都是什么人| 低密度脂蛋白高吃什么药| 纵是什么意思| 吃什么排气最快| 扬州有什么好玩的地方| 什么食物含锌最多| 软肋骨炎吃什么药对症| 苹果和生姜煮水喝有什么功效| 什么是体制内的工作| 难于上青天是什么意思| 不超过是什么意思| 偷什么东西不犯法| 闭门思过是什么意思| 宋江是属什么生肖| 脂溢性脱发是什么意思| 什么零食热量低有利于减肥| 死了妻子的男人叫什么| 夏天煲什么汤最好| vg是什么意思| 青海有什么特产| 儿童贫血吃什么补血最快| 胃食管反流能吃什么水果| 中老年补钙吃什么钙片好| 暗物质和暗能量是什么| 可可粉是什么| 盎司是什么单位| 料酒和黄酒有什么区别| 整天想睡觉是什么原因| 新生儿超敏c反应蛋白高说明什么| 番茄和西红柿有什么区别| 肠炎有什么症状表现| 脚转筋是什么原因| 鼻窦炎吃什么药| 男性内分泌科检查什么| 格色是什么意思| 二尖瓣反流吃什么药| 敌敌畏是什么| 骨质增生挂什么科| 肺上有结节是什么病| 胰岛素过高会导致什么| 三七粉主治什么病| 男人遗精是什么原因造成的| 胆囊壁毛糙吃什么药效果好| 含蓄是什么意思| 身披枷锁是什么生肖| 结膜炎角膜炎用什么眼药水| 宝宝尿少是什么原因| 丑小鸭告诉我们一个什么道理| 希五行属什么| 慢性炎伴鳞化是什么意思| 哪吒是一个什么样的人| 电导率是什么意思| 土阜念什么| 为什么禁止克隆人| 农历六月初六是什么节| 马镫什么时候发明的| 纯钛对人体有什么好处| 50是什么意思| 金国是现在的什么地方| 偏头痛什么原因引起的| a血型和o血型生出宝宝是什么血型| 723是什么意思| 梦见老虎是什么意思| va是什么意思| 橘白猫是什么品种| 什么叫戈壁滩| aj是什么| 嘴唇下面长痘痘是什么原因| 周围神经病是什么意思| 手抓饼里面夹什么好吃| 嘴唇紫色是什么原因| 7月1号是什么节| 什么的国王| 下午四点到五点是什么时辰| 桥本甲状腺炎是什么意思| 粉瘤是什么东西| 环移位了有什么症状| 手术后能吃什么| 上火吃什么药最有效果| 天麻不能和什么一起吃| 眩晕是什么症状| 习俗是什么意思| 三焦指的是什么| 双侧输尿管不扩张是什么意思| 一什么花瓶| 制动是什么意思| 后壁和前壁有什么区别| 天为什么会下雨| 离婚要带什么| 中国国花是什么花| 肠胃不舒服吃什么药| 碧霄是什么意思| 欣什么若什么| ca是什么意思| 为什么二楼比三楼好| 辩证思维是什么意思| 桃子有什么功效| 吃什么补锌| 农历10月是什么星座| 阴唇痒用什么药| 高血压能吃什么| 本是什么生肖| 奶油是什么做的| 五月掉床有什么说法| 减肥期间可以吃什么零食| 艺术有什么用| 一醉方休什么意思| 司空见惯是说司空见惯了什么| 草字头一个见念什么| 肚脐上三指是什么地方| 官员出狱后靠什么生活| 裙带菜是什么菜| 血糖高不能吃什么食物| 孩子咬嘴唇是什么原因| 什么是胰岛素抵抗| 花椒什么时候成熟| 氮肥是什么肥| 梦见和老公结婚是什么意思| 桑黄是什么东西| 前列腺炎有什么征兆| 射进去有什么感觉| 静若幽兰什么意思| rng是什么意思| 晚上睡觉口干是什么原因| 夏天感冒吃什么药| 肌肉僵硬是什么原因| street是什么意思| 什么手机拍照效果最好| 小巧玲珑是什么意思| 为什么要拔智齿| 人乳头瘤病毒是什么意思| 桂花代表什么生肖| 头晕去医院看什么科| 舒字五行属什么的| 人参和什么泡酒能壮阳| 疣是一种什么病| 儿童哮喘挂什么科| 白色代表什么| 十一月底是什么星座| 外交部发言人什么级别| 手脚发热吃什么药| 动脉瘤是什么| 什么品牌蓝牙耳机好| 中的反义词是什么| 前列腺增生吃什么药最好| 为什么会打喷嚏| 卵黄囊偏大是什么原因| 桂圆是什么| 诡辩是什么意思| 大材小用是什么生肖| 腰扭伤吃什么药最有效| 斯沃琪手表什么档次| 战战兢兢的意思是什么| 八月七号是什么星座| 白羊属于什么象星座| 三候是什么意思| 高筋面粉适合做什么| 乙酰氨基葡萄糖苷酶阳性什么意思| 石女是什么样子的| 宜子痣是什么意思| 胆汁为什么会反流到胃里面| 容易饿是什么原因| 不锈钢肥皂是什么原理| 弈字五行属什么| 92年是什么生肖| 脾胃湿热吃什么中成药| 肝做什么检查最准确| 身上遇热就痒是什么病| shia是什么意思| 汤伤用什么药| 凤毛麟角是什么生肖| 梦见摘枣是什么意思| 吃菠萝蜜有什么好处| 滑膜炎是什么病| 糖尿病人能吃什么水果| 百度

看嘴唇就能知道…

百度 本决定自2014年1月1日起施行。

In statistics, a categorical variable (also called qualitative variable) is a variable that can take on one of a limited, and usually fixed, number of possible values, assigning each individual or other unit of observation to a particular group or nominal category on the basis of some qualitative property.[1] In computer science and some branches of mathematics, categorical variables are referred to as enumerations or enumerated types. Commonly (though not in this article), each of the possible values of a categorical variable is referred to as a level. The probability distribution associated with a random categorical variable is called a categorical distribution.

Categorical data is the statistical data type consisting of categorical variables or of data that has been converted into that form, for example as grouped data. More specifically, categorical data may derive from observations made of qualitative data that are summarised as counts or cross tabulations, or from observations of quantitative data grouped within given intervals. Often, purely categorical data are summarised in the form of a contingency table. However, particularly when considering data analysis, it is common to use the term "categorical data" to apply to data sets that, while containing some categorical variables, may also contain non-categorical variables. Ordinal variables have a meaningful ordering, while nominal variables have no meaningful ordering.

A categorical variable that can take on exactly two values is termed a binary variable or a dichotomous variable; an important special case is the Bernoulli variable. Categorical variables with more than two possible values are called polytomous variables; categorical variables are often assumed to be polytomous unless otherwise specified. Discretization is treating continuous data as if it were categorical. Dichotomization is treating continuous data or polytomous variables as if they were binary variables. Regression analysis often treats category membership with one or more quantitative dummy variables.

Examples of categorical variables

edit

Examples of values that might be represented in a categorical variable:

  • Demographic information of a population: gender, disease status.
  • The blood type of a person: A, B, AB or O.
  • The political party that a voter might vote for, e.?g. Green Party, Christian Democrat, Social Democrat, etc.
  • The type of a rock: igneous, sedimentary or metamorphic.
  • The identity of a particular word (e.g., in a language model): One of V possible choices, for a vocabulary of size V.

Notation

edit

For ease in statistical processing, categorical variables may be assigned numeric indices, e.g. 1 through K for a K-way categorical variable (i.e. a variable that can express exactly K possible values). In general, however, the numbers are arbitrary, and have no significance beyond simply providing a convenient label for a particular value. In other words, the values in a categorical variable exist on a nominal scale: they each represent a logically separate concept, cannot necessarily be meaningfully ordered, and cannot be otherwise manipulated as numbers could be. Instead, valid operations are equivalence, set membership, and other set-related operations.

As a result, the central tendency of a set of categorical variables is given by its mode; neither the mean nor the median can be defined. As an example, given a set of people, we can consider the set of categorical variables corresponding to their last names. We can consider operations such as equivalence (whether two people have the same last name), set membership (whether a person has a name in a given list), counting (how many people have a given last name), or finding the mode (which name occurs most often). However, we cannot meaningfully compute the "sum" of Smith + Johnson, or ask whether Smith is "less than" or "greater than" Johnson. As a result, we cannot meaningfully ask what the "average name" (the mean) or the "middle-most name" (the median) is in a set of names.

This ignores the concept of alphabetical order, which is a property that is not inherent in the names themselves, but in the way we construct the labels. For example, if we write the names in Cyrillic and consider the Cyrillic ordering of letters, we might get a different result of evaluating "Smith < Johnson" than if we write the names in the standard Latin alphabet; and if we write the names in Chinese characters, we cannot meaningfully evaluate "Smith < Johnson" at all, because no consistent ordering is defined for such characters. However, if we do consider the names as written, e.g., in the Latin alphabet, and define an ordering corresponding to standard alphabetical order, then we have effectively converted them into ordinal variables defined on an ordinal scale.

Number of possible values

edit

Categorical random variables are normally described statistically by a categorical distribution, which allows an arbitrary K-way categorical variable to be expressed with separate probabilities specified for each of the K possible outcomes. Such multiple-category categorical variables are often analyzed using a multinomial distribution, which counts the frequency of each possible combination of numbers of occurrences of the various categories. Regression analysis on categorical outcomes is accomplished through multinomial logistic regression, multinomial probit or a related type of discrete choice model.

Categorical variables that have only two possible outcomes (e.g., "yes" vs. "no" or "success" vs. "failure") are known as binary variables (or Bernoulli variables). Because of their importance, these variables are often considered a separate category, with a separate distribution (the Bernoulli distribution) and separate regression models (logistic regression, probit regression, etc.). As a result, the term "categorical variable" is often reserved for cases with 3 or more outcomes, sometimes termed a multi-way variable in opposition to a binary variable.

It is also possible to consider categorical variables where the number of categories is not fixed in advance. As an example, for a categorical variable describing a particular word, we might not know in advance the size of the vocabulary, and we would like to allow for the possibility of encountering words that we have not already seen. Standard statistical models, such as those involving the categorical distribution and multinomial logistic regression, assume that the number of categories is known in advance, and changing the number of categories on the fly is tricky. In such cases, more advanced techniques must be used. An example is the Dirichlet process, which falls in the realm of nonparametric statistics. In such a case, it is logically assumed that an infinite number of categories exist, but at any one time most of them (in fact, all but a finite number) have never been seen. All formulas are phrased in terms of the number of categories actually seen so far rather than the (infinite) total number of potential categories in existence, and methods are created for incremental updating of statistical distributions, including adding "new" categories.

Categorical variables and regression

edit

Categorical variables represent a qualitative method of scoring data (i.e. represents categories or group membership). These can be included as independent variables in a regression analysis or as dependent variables in logistic regression or probit regression, but must be converted to quantitative data in order to be able to analyze the data. One does so through the use of coding systems. Analyses are conducted such that only g -1 (g being the number of groups) are coded. This minimizes redundancy while still representing the complete data set as no additional information would be gained from coding the total g groups: for example, when coding gender (where g = 2: male and female), if we only code females everyone left over would necessarily be males. In general, the group that one does not code for is the group of least interest.[2]

There are three main coding systems typically used in the analysis of categorical variables in regression: dummy coding, effects coding, and contrast coding. The regression equation takes the form of Y = bX + a, where b is the slope and gives the weight empirically assigned to an explanator, X is the explanatory variable, and a is the Y-intercept, and these values take on different meanings based on the coding system used. The choice of coding system does not affect the F or R2 statistics. However, one chooses a coding system based on the comparison of interest since the interpretation of b values will vary.[2]

Dummy coding

edit

Dummy coding is used when there is a control or comparison group in mind. One is therefore analyzing the data of one group in relation to the comparison group: a represents the mean of the control group and b is the difference between the mean of the experimental group and the mean of the control group. It is suggested that three criteria be met for specifying a suitable control group: the group should be a well-established group (e.g. should not be an "other" category), there should be a logical reason for selecting this group as a comparison (e.g. the group is anticipated to score highest on the dependent variable), and finally, the group's sample size should be substantive and not small compared to the other groups.[3]

In dummy coding, the reference group is assigned a value of 0 for each code variable, the group of interest for comparison to the reference group is assigned a value of 1 for its specified code variable, while all other groups are assigned 0 for that particular code variable.[2]

The b values should be interpreted such that the experimental group is being compared against the control group. Therefore, yielding a negative b value would entail the experimental group have scored less than the control group on the dependent variable. To illustrate this, suppose that we are measuring optimism among several nationalities and we have decided that French people would serve as a useful control. If we are comparing them against Italians, and we observe a negative b value, this would suggest Italians obtain lower optimism scores on average.

The following table is an example of dummy coding with French as the control group and C1, C2, and C3 respectively being the codes for Italian, German, and Other (neither French nor Italian nor German):

Nationality C1 C2 C3
French 0 0 0
Italian 1 0 0
German 0 1 0
Other 0 0 1

Effects coding

edit

In the effects coding system, data are analyzed through comparing one group to all other groups. Unlike dummy coding, there is no control group. Rather, the comparison is being made at the mean of all groups combined (a is now the grand mean). Therefore, one is not looking for data in relation to another group but rather, one is seeking data in relation to the grand mean.[2]

Effects coding can either be weighted or unweighted. Weighted effects coding is simply calculating a weighted grand mean, thus taking into account the sample size in each variable. This is most appropriate in situations where the sample is representative of the population in question. Unweighted effects coding is most appropriate in situations where differences in sample size are the result of incidental factors. The interpretation of b is different for each: in unweighted effects coding b is the difference between the mean of the experimental group and the grand mean, whereas in the weighted situation it is the mean of the experimental group minus the weighted grand mean.[2]

In effects coding, we code the group of interest with a 1, just as we would for dummy coding. The principal difference is that we code ?1 for the group we are least interested in. Since we continue to use a g - 1 coding scheme, it is in fact the ?1 coded group that will not produce data, hence the fact that we are least interested in that group. A code of 0 is assigned to all other groups.

The b values should be interpreted such that the experimental group is being compared against the mean of all groups combined (or weighted grand mean in the case of weighted effects coding). Therefore, yielding a negative b value would entail the coded group as having scored less than the mean of all groups on the dependent variable. Using our previous example of optimism scores among nationalities, if the group of interest is Italians, observing a negative b value suggest they obtain a lower optimism score.

The following table is an example of effects coding with Other as the group of least interest.

Nationality C1 C2 C3
French 0 0 1
Italian 1 0 0
German 0 1 0
Other ?1 ?1 ?1

Contrast coding

edit

The contrast coding system allows a researcher to directly ask specific questions. Rather than having the coding system dictate the comparison being made (i.e., against a control group as in dummy coding, or against all groups as in effects coding) one can design a unique comparison catering to one's specific research question. This tailored hypothesis is generally based on previous theory and/or research. The hypotheses proposed are generally as follows: first, there is the central hypothesis which postulates a large difference between two sets of groups; the second hypothesis suggests that within each set, the differences among the groups are small. Through its a priori focused hypotheses, contrast coding may yield an increase in power of the statistical test when compared with the less directed previous coding systems.[2]

Certain differences emerge when we compare our a priori coefficients between ANOVA and regression. Unlike when used in ANOVA, where it is at the researcher's discretion whether they choose coefficient values that are either orthogonal or non-orthogonal, in regression, it is essential that the coefficient values assigned in contrast coding be orthogonal. Furthermore, in regression, coefficient values must be either in fractional or decimal form. They cannot take on interval values.

The construction of contrast codes is restricted by three rules:

  1. The sum of the contrast coefficients per each code variable must equal zero.
  2. The difference between the sum of the positive coefficients and the sum of the negative coefficients should equal 1.
  3. Coded variables should be orthogonal.[2]

Violating rule 2 produces accurate R2 and F values, indicating that we would reach the same conclusions about whether or not there is a significant difference; however, we can no longer interpret the b values as a mean difference.

To illustrate the construction of contrast codes consider the following table. Coefficients were chosen to illustrate our a priori hypotheses: Hypothesis 1: French and Italian persons will score higher on optimism than Germans (French = +0.33, Italian = +0.33, German = ?0.66). This is illustrated through assigning the same coefficient to the French and Italian categories and a different one to the Germans. The signs assigned indicate the direction of the relationship (hence giving Germans a negative sign is indicative of their lower hypothesized optimism scores). Hypothesis 2: French and Italians are expected to differ on their optimism scores (French = +0.50, Italian = ?0.50, German = 0). Here, assigning a zero value to Germans demonstrates their non-inclusion in the analysis of this hypothesis. Again, the signs assigned are indicative of the proposed relationship.

Nationality C1 C2
French +0.33 +0.50
Italian +0.33 ?0.50
German ?0.66 0

Nonsense coding

edit

Nonsense coding occurs when one uses arbitrary values in place of the designated "0"s "1"s and "-1"s seen in the previous coding systems. Although it produces correct mean values for the variables, the use of nonsense coding is not recommended as it will lead to uninterpretable statistical results.[2]

Embeddings

edit

Embeddings are codings of categorical values into low-dimensional real-valued (sometimes complex-valued) vector spaces, usually in such a way that ‘similar’ values are assigned ‘similar’ vectors, or with respect to some other kind of criterion making the vectors useful for the respective application. A common special case are word embeddings, where the possible values of the categorical variable are the words in a language and words with similar meanings are to be assigned similar vectors.

Interactions

edit

An interaction may arise when considering the relationship among three or more variables, and describes a situation in which the simultaneous influence of two variables on a third is not additive. Interactions may arise with categorical variables in two ways: either categorical by categorical variable interactions, or categorical by continuous variable interactions.

Categorical by categorical variable interactions

edit

This type of interaction arises when we have two categorical variables. In order to probe this type of interaction, one would code using the system that addresses the researcher's hypothesis most appropriately. The product of the codes yields the interaction. One may then calculate the b value and determine whether the interaction is significant.[2]

Categorical by continuous variable interactions

edit

Simple slopes analysis is a common post hoc test used in regression which is similar to the simple effects analysis in ANOVA, used to analyze interactions. In this test, we are examining the simple slopes of one independent variable at specific values of the other independent variable. Such a test is not limited to use with continuous variables, but may also be employed when the independent variable is categorical. We cannot simply choose values to probe the interaction as we would in the continuous variable case because of the nominal nature of the data (i.e., in the continuous case, one could analyze the data at high, moderate, and low levels assigning 1 standard deviation above the mean, at the mean, and at one standard deviation below the mean respectively). In our categorical case we would use a simple regression equation for each group to investigate the simple slopes. It is common practice to standardize or center variables to make the data more interpretable in simple slopes analysis; however, categorical variables should never be standardized or centered. This test can be used with all coding systems.[2]

See also

edit

References

edit
  1. ^ Yates, Daniel S.; Moore, David S.; Starnes, Daren S. (2003). The Practice of Statistics (2nd?ed.). New York: Freeman. ISBN?978-0-7167-4773-4. Archived from the original on 2025-08-14. Retrieved 2025-08-14.
  2. ^ a b c d e f g h i j Cohen, J.; Cohen, P.; West, S. G.; Aiken, L. S. (2003). Applied multiple regression/correlation analysis for the behavioural sciences (3rd ed.). New York, NY: Routledge.
  3. ^ Hardy, Melissa (1993). Regression with dummy variables. Newbury Park, CA: Sage.

Further reading

edit
星期三左眼皮跳是什么预兆 因材施教什么意思 待字闺中什么意思 维生素b族有什么用 阴性什么意思
聚乙烯醇是什么材料 流产可以吃什么水果 鼓包是什么意思 送礼送什么 喝牛奶放屁多是什么原因
出现幻觉幻听是什么心理疾病 谢娜人气为什么那么高 圭是什么意思 服中药期间忌吃什么 频繁什么意思
黄酒是什么酒 人格魅力什么意思 布丁是用什么做的 糖尿病病人吃什么水果 冠脉ct和冠脉造影有什么区别
古惑仔是什么hcv8jop2ns5r.cn 大土土什么字hcv9jop5ns8r.cn 什么帽子不能戴hcv8jop1ns8r.cn 蜂蜜什么时候吃最好hcv8jop3ns3r.cn 什么是公共场所xinmaowt.com
大公鸡衣服是什么牌子hcv9jop6ns5r.cn 氧气是什么hcv7jop6ns8r.cn 头发为什么会变白hcv8jop8ns1r.cn 电磁波是什么dayuxmw.com 儿童风寒咳嗽吃什么药hcv9jop6ns2r.cn
甘油三酯偏高吃什么药hcv8jop3ns5r.cn 晚上睡觉脚抽搐是什么原因xinmaowt.com 小猫来家里有什么预兆hcv7jop5ns5r.cn 掌心痣代表什么意思hcv9jop1ns2r.cn 距离感是什么意思hcv8jop6ns8r.cn
口蘑是什么蘑菇hcv8jop2ns3r.cn 中国地图像什么hcv7jop6ns9r.cn 傍大款是什么意思hcv9jop4ns8r.cn 吃什么可以补铁helloaicloud.com 银杏树叶子像什么hcv9jop2ns7r.cn
百度