Thursday, September 26, 2019

Multiple Regression Empirical Project Essay Example | Topics and Well Written Essays - 2500 words

Multiple Regression Empirical Project - Essay Example A billion dollar increase in net exports holding consumption and direct foreign investments constant leads to 0.47 billion dollar increase in GDP. Considering consumption alone, it was found out that a billion dollar increase in consumption leads to 1.43 billion dollars increase in GDP and that consumption levels explains 99.84% of the total variations in GDP [r2 (60) = 0.9984]. Further, taking foreign direct investments alone, it was found that a billion dollar increase in foreign direct investment leads to 5.33 billion dollars increase in GDP. This model was found to be significant at 5% level of significance and that FDI explains 96.39% of the total variations in GDP Lastly, a billion dollar increase in net exports led to 17.47 billion dollars decrease in GDP and the model with NE alone was found to be significant at 5% level of significance and that NE explains 54.41% of the total variations in GDP. This study aimed at determining the impact of responsible consumption, foreign direct investments and net exports and employed the use of secondary data to proof the objectives. Different writers have argued that consumptions and investments are the key variables on which the GDP depends most. However, other variables like irresponsible consumptions, political un-rests, environmental degradation, and lack of government priorities translate to irresponsible spending are some other factors which should be taken care of for GDP to grow. GDP is the cumulative amount of goods/services which a country produces within a given year (Hall and Mishkin 1982; Hill 1992). When GDP changes, then a country is said to have experienced economic growth (if positive change) and economic melt-down (if negative growth-previous year’s performance is better than current year’s). Factors like level of consumption, direct foreign investments and net exports are some of the factors which contribute to positive GDP growth, hence economic growth (Haron 2005). High direct foreign

Speech Essay Example | Topics and Well Written Essays - 4250 words

Speech - Essay Example The three of you spend the first several minutes discussing what you did over the weekend. You all laugh over the exploits of classmate #1, who spent all night Friday and Saturday at several nightclubs. You and classmate #2 take your work out and place it on the table. Classmate #1 does not have any work completed and explains that with finals and the beginning of a head-cold, there just was not enough time. You get angry and frustrated; after all, you spent all weekend in your dorm working on this project. You look across the table at classmate #1 who is sitting with arms crossed, glaring at you. I want Classmate #1 to understand that he has not been fair, having spent the weekend nights at clubs while I and Classmate #2 had to forego a lot of things just to be able to work on the group project. After all, the project is supposed to be the result of all our efforts. The three of us would be getting equal credit based on the overall quality of the project that we will submit. I will convey my displeasure by knocking some sense into his head. I will not be affected by his defiant stance and I will go ahead to tell him that he has no right to glare at me with his arms crossed for he is the one who has wronged me and Classmate #2. I will talk in a calm but firm manner, and will unflinchingly look at Classmate #1 straight in the eye. This way, I will show him that I am not intimidated by his purported stance. I will further drive home the point that he has to shape up and make up for his lack of output so far by having the biggest share of work to be done before Wednesday. B. Using the example above, describe each component of the Speech Communication Process. a. The speaker is a college student who is enrolled in a speech class. Filled with anger and frustration about the receiver's irresponsible ways, the speaker earnestly tries to make the receiver comprehend the message that he wants to bring across. b. The receiver is one of the speaker's classmates in the speech class. Together with yet another classmate, the three of them formed a group for the completion of a certain project due in class. The receiver is the co-member who seemed to have no intention of contributing anything to their group project. Instead of being apologetic, this receiver takes on a defiant stance when he sees the speaker's reaction to his having done no part of his individual work during the weekend. c. The message is centered on their urgent need to finish their group project before the deadline. It pertains to the prevailing situation of the three-member group. The speaker means to convey the importance of fairness and of getting credit only when deserved. d. The channels include the speaker's initially angry and frustrated reaction, and his calm and firm manners of the speaker. It also covers the clarity of the speaker's message. 2. How can you control you

Wednesday, September 25, 2019

Death Penalty Essay Example | Topics and Well Written Essays - 2000 words - 1

Death Penalty - Essay Example Our main points for the abolition of the death penalty are morality and technicality. The original arguments for capital punishment does not anymore apply but are outdated – deterrence, retribution, etc. The teachings of the Bible – the Old and the New Testaments – tell us one important aspect of creation: protect life and do not allow vengeance. God did not kill Cain for killing his own brother Abel but instead sent him on exile and put a mark on him so that no one would kill him. A passage in the Bible of states: ‘an eye for an eye, a tooth for a tooth’. This teaching does not mean taking the life of a murderer or someone who has committed a heinous crime, but it means limiting the retribution for an offense. When Jesus was presented the woman accused of adultery, he did not condemn the woman but told those present â€Å"to cast the first stone†, which means we should not condemn anyone but allow a sinner to reform. Another argument is the ground of technicality. The criminal justice system, a human system, is flawed. I mean, it is not always perfect despite all the bright, legal minds the world has ever produced. Rulings are not perfect. The Supreme Court ruled that the death penalty violated the Eight and Fourteenth Amendments. Then in another ruling in 1976, Gregg v. Georgia, the Court again contradicted itself by ruling that the death penalty per se was not unconstitutional. (Bedau 2005, p. 23) In Furman, it was ruled that some state statutes were unconstitutional, which allowed that death penalty statutes had to be rewritten. Advocates of the death penalty began proposing new laws for capital punishment. In other words, advocates of the death penalty interpreted this as an opportunity to write new laws so that there would be no more doubts of retaining the death penalty. It was reported that there were 35 states that rewrote their

Tuesday, September 24, 2019

Reproductive Rights Essay Example | Topics and Well Written Essays - 250 words

Reproductive Rights - Essay Example II. Identification and Evaluation of Ethical Principles of Reproductive Rights Reproductive rights are controversial for several reasons. According to Bellieni and Buonocore (2006), some ethical principles which can be evaluated include: the potential abuse of womens’ bodies in a male-dominated medical profession; the debate over the validity of the embryo being seen as a ‘person,’ with regard to state and federal law; and the fact that studies have shown that in vitro fertilization has shown higher risks of birth defects such as cerebral palsy in children formed as a result of the procedure (pp. 93). Abortion is legal in the U.S., according to federal law. III. The Application of Principles to Ethical Issues with Various Implications One of the main arguments that naysayers usually make with regard to embryonic procedures is that embryos are actually people, and that scientists are ‘playing God’ by creating children in a scientific fashion—som etimes discarding embryos that are damaged, which to some people is unethical because a person’s life is being discontinued.

Monday, September 23, 2019

Media industry Essay Example | Topics and Well Written Essays - 1500 words

Media industry - Essay Example (Ross 2004). The situation is not very different in the advertising industry, as we will see in the discussion on the industry. Women in British Media: Many media commentators described the 1990s as the decade of women. The (then) Conservative government in Britain initiated affirmative action to employ women in large numbers in all sectors including the media industry. There was an apprehension that there would not be enough female graduates to fill job demand. In 1990, three women were actually appointed as editors of national newspapers. This is not to say that the atmosphere of gender discrimination has totally changed. Soft Jobs and the Fair Sex: Even in the allocation of routine jobs, there is a gender divide. By convention men journalists covered news stories concerning politics, crime, finance, education and upbringing; women covered 'human interest', consumer news, culture and social policy. Men covered the 'facts', 'sensation' and 'male' angles while women covered the 'background and effects', 'compassion', 'general' and 'female' angles. The sources for men were men and women were women. The ethics behind men's stories were detached while those of women's were based on social needs. (Carter 1998 36 - Adapted from table) Is it different in the U.S. ... Broadcast journalism employs 26% women as local TV news directors, 17% as local TV managers and 13% as general managers in radio stations. The silver lining of the situation according to Byerly is that many male journalists today identify with feminism and "indeed, have taken their feminist colleagues' side to protest sexist news coverage as well as to demand more egalitarian newsroom policies." (Byerly 2004 113) And Internationally : The situation is not much different internationally. A survey by the International Federation of Journalists (IFJ), world's largest journalists' association, shows that the number of women journalists varies widely: from 6% in Sri Lanka to around 50% in Finland. But at higher - decision making - levels like editors, department heads and media owners, the numbers are very low and account for barely 6%. The percentage is highest (10-20) in Cyprus, Costa Rica, Mexico and Sweden. (Byerly 2004 109-126) Social Milieu and Women in Media: When we seek to analyze the levels at which women are employed by the media industry, it is important to consider the socio-political milieu in which the industry functions and the power the industry wields. This is because media too is a product of the socio-political milieu and cannot be considered in isolation. The social power of the communication (media) industries has long been recognised "as being at the heart of contemporary societies industrially, economically, politically and culturally" (Marris et. al. (eds.) 1999 13). Electronic age and the advent of the internet have added a new powerful thrust to traditional media, which primarily consisted of print media, radio and motion pictures and

Saturday, September 21, 2019

The significance of the post-war settlement Essay Example for Free

The significance of the post-war settlement Essay After World War 2, the extent to which participating countries had lost acted as an eye opener. What these countries expected never came to be. Having fought in war 1, and then war 2, most resources were running towards depletion and worst of all masses of people lost lives, economies, weakened, infrastructure destroyed and industrial development totally demerited. Post war era was the start to a new world order mainly to be characterized by peace, potency, mutuality and prosperity. After World War 2, every country that had taken part in it was left on a stand still. Having invested most of the resources in the war that turned out to be fruitless, it was a turning point for a majority of them. Governments had to work out new sources and systems of accumulating capital, re-institutional arrangements. The period around and after World War 2 led to advancement and fundamental industrial relations re-birth. The development at that time reflected positive future stability and durability. In the whole continent, only Sweden and Switzerland experienced tranquility. All the other countries were faced with war, colonization and enmity. Post war reconstruction of national boundaries, economical atmosphere and political systems facilitated unprecedented significance in development. The change in economical, institutional perception and perspectives is a niche definition of my understanding of post war settlement. (Fulcher,1987). In my discussion, I will scrutinize two global economical giant countries, which are Germany and Russia. To start, after World War 2 life was indeed very tough in the Federal Republic of Germany, with almost all systems in a shattered condition. Very hard living standards came up as a result of food shortage, diseases, instability, lack of job opportunities and unemployment. As a result activism and demonstrations took center stage, this was made possible by the Germany’s unionification. All employees were under the same confederation, which pushed for further reforms putting better living standards in their priority. This move turned out successive due to the fact that employees and unions existed on a mutually potent relationship, that is, one could not live without the other. Citizen thought that separation of politics and industrial issues would be a merit to their welfare. Another development that followed in turn was economical stability and expansion. To boost their progress, Britain sent food relief in 1946. Later on, Germany saw itself join membership of EEC in 1957 through the help of a strategic plan known as the `Marshal Plan’. Old industrious organizations, firms and companies remain stable as growth began to be felt. Post wartime is seen as a stability period in most countries that experienced the effects of war- Volkswagen, the automobile manufacturer is one of the companies that lived on and continue to today. Germany was forced to concentrate all it effort in policies and strategies towards economic growth. They had to halt active political presence for decades. With serious considerations being put into practice, Germany woke up to an economic excellence. As an advantage of post war activity, Germany became, and still is, among the giant economical strengths in Europe and universally. In a bid to bounce back, post wartime witnessed the practice of mass production, industries embarked on manufacturing of goods in surplus. As a result it attracted German citizens to mass consumption. This was a great move since the more they produced the more it was consumed. The gross net profit grew drastically so did their economy. Life felt cozier as job opportunities increased because of the mushrooming of many industries. Politically the country was shaped newly completely for quite a long time, at one time Germany had to lay low. They had been toppled completely and their Nazi regime wrecked down. This turn of events saw its leaders tried at Nuremberg for crimes against human rights, they had to face justice in their own home place or rather their site of propaganda brilliance. (Skidelsky, 1979) Although the late tyrant leader Adolph Hitler escaped trial and execution at that time: he later committed suicide in Berlin, at the end of the war. He felt so intimidated of the counts he would be changed for. World War 2 left many German capital towns in ruins from the massive bombings carried out in it. Germany was segregated into zones by capability and powers; this in turn resulted to a permanent political stability and settlement. The European Union that strongly stands out today has it motherly roots from world war 2 time. In fact it grew from the European Coal and Steel Community (ECSC), Its primary intention was to develop the steel and coal resources from member states, support and boost their economies.the ECSC facilitated the diffusion of the tensions that had resulted between enemy countries in world war 2. With time, this economic co-operation/merger grew enoumously , adding in new members hence broadening their scope, later the European Union was formed from EEC, European Economic Community. Many other prominent organizations today have their history date back to World War 2; for example, the world trade organization,WTO, the United Nations, the international monetary fund IMF, the World Bank in which then West Germany but now Germany had taken part in its formation and stand as members. Another very important significance is the evolution of initial and the follow up of advanced technological progress that had captivated interest during the war. The development was reflected in almost all industrial fields: in electronics, computers. These advancements helped Germany create the foundation for its realization into further development. This finally transformed to what was referred to as the post world war 2 world. The new technology, assisted in the efforts to fight diseases that had erupted during the war. Massive research, monitoring, evaluation and development, quickly attained nuclear power which had adverse impacts on the scientific fraternity, creating a network of laboratories in the whole country therafter . In addition, the struggling to crack codes initiated computer technology. There were social effects too which significantly changed almost all war participants to new degrees. One of them was increased involvement of women in the working force which replaced the place taken by many men during the war: though this initiative was reduced in the following years, due to the fast changing society. The system forced women to taking care of home and family oriented matters.(Rowthorn en.others,1992). In military aspects, the World War II had established the beginning of airpower. Germany could not be left out, this was an opportunity to draft active self defence system. They highly concentrated on project they had earlier established. Advanced aircraft composing of aeroplanes, jet fighters and missiles developed earlier saw further developments. Battle ships and tanks sprout out into the ever growing competition, with guns and ammunination reaching to leathal dynamics. Air power capability is today a constituent in any military operation or mission. (Seidman,1950) Russia on the other hand is said to be the main beneficiary when border revision was done. It saw Poland, Finland, Romania and other countries pushed into Russian boundery in their favor. Germany was not considered in the process. As a result their land area became larger creating room for development. Infrastractural systems remain almost the way they used to in Russia. Only few places had their system completely damaged. Compared to Germany, this was better off since repairs would cost less. Post war saw Germany’s roads and communicational networks left completely in a shuter. Though very many people lost lives during world war 2, in Russia than in Germany, at some point their demograpfic figure took a new turn. This was alos facilitated by increased agricultural activity. Most of Russia’s citizens depended on agriculture, fishing, forestry, or craftworks. Due to increment on the agicultural industry, their economical strength also took a stable ground-this is a major similarity that was experienced in almost all countries that had experienced war. (Gourevitch ,1985) Immortatility rates decreased and in turn working population had a future expectation. Hunger was kicked out and fertility saw the countries graphs rise demeritically. Having food and medicine eradicated the diseases that had become threatening. Politically, Russians remained constantly active compared to Germans who had to gio slow for decades. This was a crafty thing to concentrate on after the war. The approach given was also great to say, it was trying to balance business and politics. Like their counterparts the Germans, unemployment at the beginning of world war 2 hit their population badly, this led to workers demonstrations all over the country. Trade unions at the time wanted drastic change to help improve their living standards. This though was not left to spread like any other bush fire, heavy police and army contigence ensured that a runing battle existed toi keep the country tranquil and the demonstrators at bay. (Taylor,1989) Russia had severe problems following their decision to turn captured prisoners of war into plantation slave labourers. This is another reason that led to activism towards fighting for humanity/human rights. Every significant effect practiced found its relation with industrial settlement. Looking at their military state, Russia developed also in terms of strategic ideas and policies. Industrial inovations led to further outstanding developments in their manufacturing industry. This move also resulted to an interest in nuclear power. In fact, Russia heavilly invested on the project. Facing challenges though, the idea had to be carried out in top secresy due to the effect that had been seen at Nagasaki and Hiroshima. (Ferner en. Others,1994 World super power countries felt empathetic of what had happen and so took it as a mandate to control nuclear power. Countries would only be allowed to use it for developmental purposes and destruction. Russia was the most highly sort after provider of nuclear energy. They made tremendous sale that in turn aided them in developing the state. It also provided job opportunities to it people. Peace settlements characterised both Germany and Russia, this positive move indicated brighter future. Since Russia was facing challenges of cold war because of several stauch stands on their accord they saw it an apportunity to reconcile with enemy states together with Germany. The negotiations indicated that the initiative was to be the last war and a new beginning to everlasting peace. They agreed at all peace summit meetings held in Paris that idealism would take over rivalism. This was further pushed by existence of the league of nations. Expectations reached as far as waterways internationalization. This was a step ahead towards cohesive industrial relations. Independence was being experience in countries that had been colonized by Germany, they had decided to let own rule and democracy prevail. In contrary though, Russia in their side continued too occupy countries like Latvia and others with iron and steel mineral deposits or resources. Though they remain under Russia these countries witnessed new traditions that in deed very good. There was self freedom for everyone, movement was very safe generally. As a result, people grew some self determination which help industrialize Russia particularly, it was all busy-busy to earn a living a condition that turned out to be very superb.( Ryder, 2008) Preparations for compensation is another significant effort that saw ist launch in Germany more than Russia, a lot appeared to be done by Germany to cater on the aspect of enmity/crisis and conflict. Russia became laxile on this one which negatively impacted on their relationship with most countries. In turn, their industrial relations and development faced difficulties and setbacks in their wake to cold war period thereafter. Property ownership was also top after world wra 2, normally public resources were subject or prone to seizure or confiscation, by the then victorious super power countries. Many industries saw ownership and income channelled to their banks but after the war, rules changed. This saw native citizens own proprerties, industries, plantations which was a brilliant sign that the world was changing. This brought an understanding between waring communties who had different and diverse cultural variance. Although Russsia as well as Germany felt like they had lost or unfortunately settled at nothing compared to first world war, They later realized that peace was even sweeter and priceless, at some point. World war 1 settlement was not perfect, that is why war two broke out but this time around, it would be the last time blood spilled. People in fact were turning closer to God at the time, in Germany for example it saw the demise or end of the rough Nazi regime, churches were established in the folloe up.Stability in capability to keep and maintain order is something that came in world war 2 post war settlement. As a result of peace prevalence all the energies were shifted to industrial innovations and development in both Germany and Russia. In conclusion, it can said that the significance of post world war 2 was almost the same in both these two countries. The extent to which states experienced the war was relative to how harder they would be forced to work in order to achieve stability. For instance, Germany had suffered majorly in most important industries plus they had their reservouir flow of capital from their colonies in Europe and other continents stop in the wake of freedom and independence. Although it was almost incapable to bounce back they had undying determination.( Theory, 2008) On Russia case is all the same thing in all industrial developments, only that they had resources and capital that made it easier on them to progress. Politically, too, they saw a major and notable change but not as in Germany whose Nazi regime end and Hitler’s death became the starting point to humanity, democracy and of course their core economical booster indistrial stability. Reference: Fulcher J(1987). `Labour Movement Theory versus Corporatism`, Sociology Vol. 21 no 2, pp 231-252 Hyman R in Hyman R. and A. Ferner(1994) New Frontiers in European IndustrialRelations Blackwell London, Chapter one Jukka Pekkarinen, Matti Pohjola, and Bob Rowthorn (1992). Social corporatism asuperior economic system? Clarendon Press330.17 Joel Seidman(1950) Industrial and Labor Relations Review, Vol. 4, pp. 55-69 Published by: Cornell University, School of Industrial Labor Relations Peter Gourevitch (1985) Unions and Economic Change: Britain, West Germany and Sweden 331.881 GOU Ryder (2008) Post war economy: retrieved from the World Wide Web at http://www.geocities.com/Athens/Rhodes/6916/ww2.htm Skidelsky R (1979) The decline of Keynesian politics: State and economy in Contemporary capitalism: Croom Helm edition, London Taylor A J (1989) Trade Unions and Politics: a comparative introduction,Basingstoke,Macmillan Theory (2008) The world since 1945: retrieved from the World Wide Web at http://www.flowofhistory.com/units/etc/19/FC128

Friday, September 20, 2019

Data Pre-processing Tool

Data Pre-processing Tool Chapter- 2 Real life data rarely comply with the necessities of various data mining tools. It is usually inconsistent and noisy. It may contain redundant attributes, unsuitable formats etc. Hence data has to be prepared vigilantly before the data mining actually starts. It is well known fact that success of a data mining algorithm is very much dependent on the quality of data processing. Data processing is one of the most important tasks in data mining. In this context it is natural that data pre-processing is a complicated task involving large data sets. Sometimes data pre-processing take more than 50% of the total time spent in solving the data mining problem. It is crucial for data miners to choose efficient data preprocessing technique for specific data set which can not only save processing time but also retain the quality of the data for data mining process. A data pre-processing tool should help miners with many data mining activates. For example, data may be provided in different formats as discussed in previous chapter (flat files, database files etc). Data files may also have different formats of values, calculation of derived attributes, data filters, joined data sets etc. Data mining process generally starts with understanding of data. In this stage pre-processing tools may help with data exploration and data discovery tasks. Data processing includes lots of tedious works, Data pre-processing generally consists of Data Cleaning Data Integration Data Transformation And Data Reduction. In this chapter we will study all these data pre-processing activities. 2.1 Data Understanding In Data understanding phase the first task is to collect initial data and then proceed with activities in order to get well known with data, to discover data quality problems, to discover first insight into the data or to identify interesting subset to form hypothesis for hidden information. The data understanding phase according to CRISP model can be shown in following . 2.1.1 Collect Initial Data The initial collection of data includes loading of data if required for data understanding. For instance, if specific tool is applied for data understanding, it makes great sense to load your data into this tool. This attempt possibly leads to initial data preparation steps. However if data is obtained from multiple data sources then integration is an additional issue. 2.1.2 Describe data Here the gross or surface properties of the gathered data are examined. 2.1.3 Explore data This task is required to handle the data mining questions, which may be addressed using querying, visualization and reporting. These include: Sharing of key attributes, for instance the goal attribute of a prediction task Relations between pairs or small numbers of attributes Results of simple aggregations Properties of important sub-populations Simple statistical analyses. 2.1.4 Verify data quality In this step quality of data is examined. It answers questions such as: Is the data complete (does it cover all the cases required)? Is it accurate or does it contains errors and if there are errors how common are they? Are there missing values in the data? If so how are they represented, where do they occur and how common are they? 2.2 Data Preprocessing Data preprocessing phase focus on the pre-processing steps that produce the data to be mined. Data preparation or preprocessing is one most important step in data mining. Industrial practice indicates that one data is well prepared; the mined results are much more accurate. This means this step is also a very critical fro success of data mining method. Among others, data preparation mainly involves data cleaning, data integration, data transformation, and reduction. 2.2.1 Data Cleaning Data cleaning is also known as data cleansing or scrubbing. It deals with detecting and removing inconsistencies and errors from data in order to get better quality data. While using a single data source such as flat files or databases data quality problems arises due to misspellings while data entry, missing information or other invalid data. While the data is taken from the integration of multiple data sources such as data warehouses, federated database systems or global web-based information systems, the requirement for data cleaning increases significantly. This is because the multiple sources may contain redundant data in different formats. Consolidation of different data formats abs elimination of redundant information becomes necessary in order to provide access to accurate and consistent data. Good quality data requires passing a set of quality criteria. Those criteria include: Accuracy: Accuracy is an aggregated value over the criteria of integrity, consistency and density. Integrity: Integrity is an aggregated value over the criteria of completeness and validity. Completeness: completeness is achieved by correcting data containing anomalies. Validity: Validity is approximated by the amount of data satisfying integrity constraints. Consistency: consistency concerns contradictions and syntactical anomalies in data. Uniformity: it is directly related to irregularities in data. Density: The density is the quotient of missing values in the data and the number of total values ought to be known. Uniqueness: uniqueness is related to the number of duplicates present in the data. 2.2.1.1 Terms Related to Data Cleaning Data cleaning: data cleaning is the process of detecting, diagnosing, and editing damaged data. Data editing: data editing means changing the value of data which are incorrect. Data flow: data flow is defined as passing of recorded information through succeeding information carriers. Inliers: Inliers are data values falling inside the projected range. Outlier: outliers are data value falling outside the projected range. Robust estimation: evaluation of statistical parameters, using methods that are less responsive to the effect of outliers than more conventional methods are called robust method. 2.2.1.2 Definition: Data Cleaning Data cleaning is a process used to identify imprecise, incomplete, or irrational data and then improving the quality through correction of detected errors and omissions. This process may include format checks Completeness checks Reasonableness checks Limit checks Review of the data to identify outliers or other errors Assessment of data by subject area experts (e.g. taxonomic specialists). By this process suspected records are flagged, documented and checked subsequently. And finally these suspected records can be corrected. Sometimes validation checks also involve checking for compliance against applicable standards, rules, and conventions. The general framework for data cleaning given as: Define and determine error types; Search and identify error instances; Correct the errors; Document error instances and error types; and Modify data entry procedures to reduce future errors. Data cleaning process is referred by different people by a number of terms. It is a matter of preference what one uses. These terms include: Error Checking, Error Detection, Data Validation, Data Cleaning, Data Cleansing, Data Scrubbing and Error Correction. We use Data Cleaning to encompass three sub-processes, viz. Data checking and error detection; Data validation; and Error correction. A fourth improvement of the error prevention processes could perhaps be added. 2.2.1.3 Problems with Data Here we just note some key problems with data Missing data : This problem occur because of two main reasons Data are absent in source where it is expected to be present. Some times data is present are not available in appropriately form Detecting missing data is usually straightforward and simpler. Erroneous data: This problem occurs when a wrong value is recorded for a real world value. Detection of erroneous data can be quite difficult. (For instance the incorrect spelling of a name) Duplicated data : This problem occur because of two reasons Repeated entry of same real world entity with some different values Some times a real world entity may have different identifications. Repeat records are regular and frequently easy to detect. The different identification of the same real world entities can be a very hard problem to identify and solve. Heterogeneities: When data from different sources are brought together in one analysis problem heterogeneity may occur. Heterogeneity could be Structural heterogeneity arises when the data structures reflect different business usage Semantic heterogeneity arises when the meaning of data is different n each system that is being combined Heterogeneities are usually very difficult to resolve since because they usually involve a lot of contextual data that is not well defined as metadata. Information dependencies in the relationship between the different sets of attribute are commonly present. Wrong cleaning mechanisms can further damage the information in the data. Various analysis tools handle these problems in different ways. Commercial offerings are available that assist the cleaning process, but these are often problem specific. Uncertainty in information systems is a well-recognized hard problem. In following a very simple examples of missing and erroneous data is shown Extensive support for data cleaning must be provided by data warehouses. Data warehouses have high probability of â€Å"dirty data† since they load and continuously refresh huge amounts of data from a variety of sources. Since these data warehouses are used for strategic decision making therefore the correctness of their data is important to avoid wrong decisions. The ETL (Extraction, Transformation, and Loading) process for building a data warehouse is illustrated in following . Data transformations are related with schema or data translation and integration, and with filtering and aggregating data to be stored in the data warehouse. All data cleaning is classically performed in a separate data performance area prior to loading the transformed data into the warehouse. A large number of tools of varying functionality are available to support these tasks, but often a significant portion of the cleaning and transformation work has to be done manually or by low-level programs that are difficult to write and maintain. A data cleaning method should assure following: It should identify and eliminate all major errors and inconsistencies in an individual data sources and also when integrating multiple sources. Data cleaning should be supported by tools to bound manual examination and programming effort and it should be extensible so that can cover additional sources. It should be performed in association with schema related data transformations based on metadata. Data cleaning mapping functions should be specified in a declarative way and be reusable for other data sources. 2.2.1.4 Data Cleaning: Phases 1. Analysis: To identify errors and inconsistencies in the database there is a need of detailed analysis, which involves both manual inspection and automated analysis programs. This reveals where (most of) the problems are present. 2. Defining Transformation and Mapping Rules: After discovering the problems, this phase are related with defining the manner by which we are going to automate the solutions to clean the data. We will find various problems that translate to a list of activities as a result of analysis phase. Example: Remove all entries for J. Smith because they are duplicates of John Smith Find entries with `bule in colour field and change these to `blue. Find all records where the Phone number field does not match the pattern (NNNNN NNNNNN). Further steps for cleaning this data are then applied. Etc †¦ 3. Verification: In this phase we check and assess the transformation plans made in phase- 2. Without this step, we may end up making the data dirtier rather than cleaner. Since data transformation is the main step that actually changes the data itself so there is a need to be sure that the applied transformations will do it correctly. Therefore test and examine the transformation plans very carefully. Example: Let we have a very thick C++ book where it says strict in all the places where it should say struct 4. Transformation: Now if it is sure that cleaning will be done correctly, then apply the transformation verified in last step. For large database, this task is supported by a variety of tools Backflow of Cleaned Data: In a data mining the main objective is to convert and move clean data into target system. This asks for a requirement to purify legacy data. Cleansing can be a complicated process depending on the technique chosen and has to be designed carefully to achieve the objective of removal of dirty data. Some methods to accomplish the task of data cleansing of legacy system include: n Automated data cleansing n Manual data cleansing n The combined cleansing process 2.2.1.5 Missing Values Data cleaning addresses a variety of data quality problems, including noise and outliers, inconsistent data, duplicate data, and missing values. Missing values is one important problem to be addressed. Missing value problem occurs because many tuples may have no record for several attributes. For Example there is a customer sales database consisting of a whole bunch of records (lets say around 100,000) where some of the records have certain fields missing. Lets say customer income in sales data may be missing. Goal here is to find a way to predict what the missing data values should be (so that these can be filled) based on the existing data. Missing data may be due to following reasons Equipment malfunction Inconsistent with other recorded data and thus deleted Data not entered due to misunderstanding Certain data may not be considered important at the time of entry Not register history or changes of the data How to Handle Missing Values? Dealing with missing values is a regular question that has to do with the actual meaning of the data. There are various methods for handling missing entries 1. Ignore the data row. One solution of missing values is to just ignore the entire data row. This is generally done when the class label is not there (here we are assuming that the data mining goal is classification), or many attributes are missing from the row (not just one). But if the percentage of such rows is high we will definitely get a poor performance. 2. Use a global constant to fill in for missing values. We can fill in a global constant for missing values such as unknown, N/A or minus infinity. This is done because at times is just doesnt make sense to try and predict the missing value. For example if in customer sales database if, say, office address is missing for some, filling it in doesnt make much sense. This method is simple but is not full proof. 3. Use attribute mean. Let say if the average income of a a family is X you can use that value to replace missing income values in the customer sales database. 4. Use attribute mean for all samples belonging to the same class. Lets say you have a cars pricing DB that, among other things, classifies cars to Luxury and Low budget and youre dealing with missing values in the cost field. Replacing missing cost of a luxury car with the average cost of all luxury cars is probably more accurate then the value youd get if you factor in the low budget 5. Use data mining algorithm to predict the value. The value can be determined using regression, inference based tools using Bayesian formalism, decision trees, clustering algorithms etc. 2.2.1.6 Noisy Data Noise can be defined as a random error or variance in a measured variable. Due to randomness it is very difficult to follow a strategy for noise removal from the data. Real world data is not always faultless. It can suffer from corruption which may impact the interpretations of the data, models created from the data, and decisions made based on the data. Incorrect attribute values could be present because of following reasons Faulty data collection instruments Data entry problems Duplicate records Incomplete data: Inconsistent data Incorrect processing Data transmission problems Technology limitation. Inconsistency in naming convention Outliers How to handle Noisy Data? The methods for removing noise from data are as follows. 1. Binning: this approach first sort data and partition it into (equal-frequency) bins then one can smooth it using- Bin means, smooth using bin median, smooth using bin boundaries, etc. 2. Regression: in this method smoothing is done by fitting the data into regression functions. 3. Clustering: clustering detect and remove outliers from the data. 4. Combined computer and human inspection: in this approach computer detects suspicious values which are then checked by human experts (e.g., this approach deal with possible outliers).. Following methods are explained in detail as follows: Binning: Data preparation activity that converts continuous data to discrete data by replacing a value from a continuous range with a bin identifier, where each bin represents a range of values. For instance, age can be changed to bins such as 20 or under, 21-40, 41-65 and over 65. Binning methods smooth a sorted data set by consulting values around it. This is therefore called local smoothing. Let consider a binning example Binning Methods n Equal-width (distance) partitioning Divides the range into N intervals of equal size: uniform grid if A and B are the lowest and highest values of the attribute, the width of intervals will be: W = (B-A)/N. The most straightforward, but outliers may dominate presentation Skewed data is not handled well n Equal-depth (frequency) partitioning 1. It divides the range (values of a given attribute) into N intervals, each containing approximately same number of samples (elements) 2. Good data scaling 3. Managing categorical attributes can be tricky. n Smooth by bin means- Each bin value is replaced by the mean of values n Smooth by bin medians- Each bin value is replaced by the median of values n Smooth by bin boundaries Each bin value is replaced by the closest boundary value Example Let Sorted data for price (in dollars): 4, 8, 9, 15, 21, 21, 24, 25, 26, 28, 29, 34 n Partition into equal-frequency (equi-depth) bins: o Bin 1: 4, 8, 9, 15 o Bin 2: 21, 21, 24, 25 o Bin 3: 26, 28, 29, 34 n Smoothing by bin means: o Bin 1: 9, 9, 9, 9 ( for example mean of 4, 8, 9, 15 is 9) o Bin 2: 23, 23, 23, 23 o Bin 3: 29, 29, 29, 29 n Smoothing by bin boundaries: o Bin 1: 4, 4, 4, 15 o Bin 2: 21, 21, 25, 25 o Bin 3: 26, 26, 26, 34 Regression: Regression is a DM technique used to fit an equation to a dataset. The simplest form of regression is linear regression which uses the formula of a straight line (y = b+ wx) and determines the suitable values for b and w to predict the value of y based upon a given value of x. Sophisticated techniques, such as multiple regression, permit the use of more than one input variable and allow for the fitting of more complex models, such as a quadratic equation. Regression is further described in subsequent chapter while discussing predictions. Clustering: clustering is a method of grouping data into different groups , so that data in each group share similar trends and patterns. Clustering constitute a major class of data mining algorithms. These algorithms automatically partitions the data space into set of regions or cluster. The goal of the process is to find all set of similar examples in data, in some optimal fashion. Following shows three clusters. Values that fall outsid e the cluster are outliers. 4. Combined computer and human inspection: These methods find the suspicious values using the computer programs and then they are verified by human experts. By this process all outliers are checked. 2.2.1.7 Data cleaning as a process Data cleaning is the process of Detecting, Diagnosing, and Editing Data. Data cleaning is a three stage method involving repeated cycle of screening, diagnosing, and editing of suspected data abnormalities. Many data errors are detected by the way during study activities. However, it is more efficient to discover inconsistencies by actively searching for them in a planned manner. It is not always right away clear whether a data point is erroneous. Many times it requires careful examination. Likewise, missing values require additional check. Therefore, predefined rules for dealing with errors and true missing and extreme values are part of good practice. One can monitor for suspect features in survey questionnaires, databases, or analysis data. In small studies, with the examiner intimately involved at all stages, there may be small or no difference between a database and an analysis dataset. During as well as after treatment, the diagnostic and treatment phases of cleaning need insight into the sources and types of errors at all stages of the study. Data flow concept is therefore crucial in this respect. After measurement the research data go through repeated steps of- entering into information carriers, extracted, and transferred to other carriers, edited, selected, transformed, summarized, and presented. It is essential to understand that errors can occur at any stage of the data flow, including during data cleaning itself. Most of these problems are due to human error. Inaccuracy of a single data point and measurement may be tolerable, and associated to the inherent technological error of the measurement device. Therefore the process of data clenaning mus focus on those errors that are beyond small technical variations and that form a major shift within or beyond the population distribution. In turn, it must be based on understanding of technical errors and expected ranges of normal values. Some errors are worthy of higher priority, but which ones are most significant is highly study-specific. For instance in most medical epidemiological studies, errors that need to be cleaned, at all costs, include missing gender, gender misspecification, birth date or examination date errors, duplications or merging of records, and biologically impossible results. Another example is in nutrition studies, date errors lead to age errors, which in turn lead to errors in weight-for-age scoring and, further, to misclassification of subjects as under- or overweight. Errors of sex and date are particularly important because they contaminate derived variables. Prioritization is essential if the study is under time pressures or if resources for data cleaning are limited. 2.2.2 Data Integration This is a process of taking data from one or more sources and mapping it, field by field, onto a new data structure. Idea is to combine data from multiple sources into a coherent form. Various data mining projects requires data from multiple sources because n Data may be distributed over different databases or data warehouses. (for example an epidemiological study that needs information about hospital admissions and car accidents) n Sometimes data may be required from different geographic distributions, or there may be need for historical data. (e.g. integrate historical data into a new data warehouse) n There may be a necessity of enhancement of data with additional (external) data. (for improving data mining precision) 2.2.2.1 Data Integration Issues There are number of issues in data integrations. Consider two database tables. Imagine two database tables Database Table-1 Database Table-2 In integration of there two tables there are variety of issues involved such as 1. The same attribute may have different names (for example in above tables Name and Given Name are same attributes with different names) 2. An attribute may be derived from another (for example attribute Age is derived from attribute DOB) 3. Attributes might be redundant( For example attribute PID is redundant) 4. Values in attributes might be different (for example for PID 4791 values in second and third field are different in both the tables) 5. Duplicate records under different keys( there is a possibility of replication of same record with different key values) Therefore schema integration and object matching can be trickier. Question here is how equivalent entities from different sources are matched? This problem is known as entity identification problem. Conflicts have to be detected and resolved. Integration becomes easier if unique entity keys are available in all the data sets (or tables) to be linked. Metadata can help in schema integration (example of metadata for each attribute includes the name, meaning, data type and range of values permitted for the attribute) 2.2.2.1 Redundancy Redundancy is another important issue in data integration. Two given attribute (such as DOB and age for instance in give table) may be redundant if one is derived form the other attribute or set of attributes. Inconsistencies in attribute or dimension naming can lead to redundancies in the given data sets. Handling Redundant Data We can handle data redundancy problems by following ways n Use correlation analysis n Different coding / representation has to be considered (e.g. metric / imperial measures) n Careful (manual) integration of the data can reduce or prevent redundancies (and inconsistencies) n De-duplication (also called internal data linkage) o If no unique entity keys are available o Analysis of values in attributes to find duplicates n Process redundant and inconsistent data (easy if values are the same) o Delete one of the values o Average values (only for numerical attributes) o Take majority values (if more than 2 duplicates and some values are the same) Correlation analysis is explained in detail here. Correlation analysis (also called Pearsons product moment coefficient): some redundancies can be detected by using correlation analysis. Given two attributes, such analysis can measure how strong one attribute implies another. For numerical attribute we can compute correlation coefficient of two attributes A and B to evaluate the correlation between them. This is given by Where n n is the number of tuples, n and are the respective means of A and B n ÏÆ'A and ÏÆ'B are the respective standard deviation of A and B n ÃŽ £(AB) is the sum of the AB cross-product. a. If -1 b. If rA, B is equal to zero it indicates A and B are independent of each other and there is no correlation between them. c. If rA, B is less than zero then A and B are negatively correlated. , where if value of one attribute increases value of another attribute decreases. This means that one attribute discourages another attribute. It is important to note that correlation does not imply causality. That is, if A and B are correlated, this does not essentially mean that A causes B or that B causes A. for example in analyzing a demographic database, we may find that attribute representing number of accidents and the number of car theft in a region are correlated. This does not mean that one is related to another. Both may be related to third attribute, namely population. For discrete data, a correlation relation between two attributes, can be discovered by a χ ²(chi-square) test. Let A has c distinct values a1,a2,†¦Ã¢â‚¬ ¦ac and B has r different values namely b1,b2,†¦Ã¢â‚¬ ¦br The data tuple described by A and B are shown as contingency table, with c values of A (making up columns) and r values of B( making up rows). Each and every (Ai, Bj) cell in table has. X^2 = sum_{i=1}^{r} sum_{j=1}^{c} {(O_{i,j} E_{i,j})^2 over E_{i,j}} . Where n Oi, j is the observed frequency (i.e. actual count) of joint event (Ai, Bj) and n Ei, j is the expected frequency which can be computed as E_{i,j}=frac{sum_{k=1}^{c} O_{i,k} sum_{k=1}^{r} O_{k,j}}{N} , , Where n N is number of data tuple n Oi,k is number of tuples having value ai for A n Ok,j is number of tuples having value bj for B The larger the χ ² value, the more likely the variables are related. The cells that contribute the most to the χ ² value are those whose actual count is very different from the expected count Chi-Square Calculation: An Example Suppose a group of 1,500 people were surveyed. The gender of each person was noted. Each person has polled their preferred type of reading material as fiction or non-fiction. The observed frequency of each possible joint event is summarized in following table.( number in parenthesis are expected frequencies) . Calculate chi square. Play chess Not play chess Sum (row) Like science fiction 250(90) 200(360) 450 Not like science fiction 50(210) 1000(840) 1050 Sum(col.) 300 1200 1500 E11 = count (male)*count(fiction)/N = 300 * 450 / 1500 =90 and so on For this table the degree of freedom are (2-1)(2-1) =1 as table is 2X2. for 1 degree of freedom , the χ ² value needed to reject the hypothesis at the 0.001 significance level is 10.828 (taken from the table of upper percentage point of the χ ² distribution typically available in any statistic text book). Since the computed value is above this, we can reject the hypothesis that gender and preferred reading are independent and conclude that two attributes are strongly correlated for given group. Duplication must also be detected at the tuple level. The use of renormalized tables is also a source of redundancies. Redundancies may further lead to data inconsistencies (due to updating some but not others). 2.2.2.2 Detection and resolution of data value conflicts Another significant issue in data integration is the discovery and resolution of data value conflicts. For example, for the same entity, attribute values from different sources may differ. For example weight can be stored in metric unit in one source and British imperial unit in another source. For instance, for a hotel cha Data Pre-processing Tool Data Pre-processing Tool Chapter- 2 Real life data rarely comply with the necessities of various data mining tools. It is usually inconsistent and noisy. It may contain redundant attributes, unsuitable formats etc. Hence data has to be prepared vigilantly before the data mining actually starts. It is well known fact that success of a data mining algorithm is very much dependent on the quality of data processing. Data processing is one of the most important tasks in data mining. In this context it is natural that data pre-processing is a complicated task involving large data sets. Sometimes data pre-processing take more than 50% of the total time spent in solving the data mining problem. It is crucial for data miners to choose efficient data preprocessing technique for specific data set which can not only save processing time but also retain the quality of the data for data mining process. A data pre-processing tool should help miners with many data mining activates. For example, data may be provided in different formats as discussed in previous chapter (flat files, database files etc). Data files may also have different formats of values, calculation of derived attributes, data filters, joined data sets etc. Data mining process generally starts with understanding of data. In this stage pre-processing tools may help with data exploration and data discovery tasks. Data processing includes lots of tedious works, Data pre-processing generally consists of Data Cleaning Data Integration Data Transformation And Data Reduction. In this chapter we will study all these data pre-processing activities. 2.1 Data Understanding In Data understanding phase the first task is to collect initial data and then proceed with activities in order to get well known with data, to discover data quality problems, to discover first insight into the data or to identify interesting subset to form hypothesis for hidden information. The data understanding phase according to CRISP model can be shown in following . 2.1.1 Collect Initial Data The initial collection of data includes loading of data if required for data understanding. For instance, if specific tool is applied for data understanding, it makes great sense to load your data into this tool. This attempt possibly leads to initial data preparation steps. However if data is obtained from multiple data sources then integration is an additional issue. 2.1.2 Describe data Here the gross or surface properties of the gathered data are examined. 2.1.3 Explore data This task is required to handle the data mining questions, which may be addressed using querying, visualization and reporting. These include: Sharing of key attributes, for instance the goal attribute of a prediction task Relations between pairs or small numbers of attributes Results of simple aggregations Properties of important sub-populations Simple statistical analyses. 2.1.4 Verify data quality In this step quality of data is examined. It answers questions such as: Is the data complete (does it cover all the cases required)? Is it accurate or does it contains errors and if there are errors how common are they? Are there missing values in the data? If so how are they represented, where do they occur and how common are they? 2.2 Data Preprocessing Data preprocessing phase focus on the pre-processing steps that produce the data to be mined. Data preparation or preprocessing is one most important step in data mining. Industrial practice indicates that one data is well prepared; the mined results are much more accurate. This means this step is also a very critical fro success of data mining method. Among others, data preparation mainly involves data cleaning, data integration, data transformation, and reduction. 2.2.1 Data Cleaning Data cleaning is also known as data cleansing or scrubbing. It deals with detecting and removing inconsistencies and errors from data in order to get better quality data. While using a single data source such as flat files or databases data quality problems arises due to misspellings while data entry, missing information or other invalid data. While the data is taken from the integration of multiple data sources such as data warehouses, federated database systems or global web-based information systems, the requirement for data cleaning increases significantly. This is because the multiple sources may contain redundant data in different formats. Consolidation of different data formats abs elimination of redundant information becomes necessary in order to provide access to accurate and consistent data. Good quality data requires passing a set of quality criteria. Those criteria include: Accuracy: Accuracy is an aggregated value over the criteria of integrity, consistency and density. Integrity: Integrity is an aggregated value over the criteria of completeness and validity. Completeness: completeness is achieved by correcting data containing anomalies. Validity: Validity is approximated by the amount of data satisfying integrity constraints. Consistency: consistency concerns contradictions and syntactical anomalies in data. Uniformity: it is directly related to irregularities in data. Density: The density is the quotient of missing values in the data and the number of total values ought to be known. Uniqueness: uniqueness is related to the number of duplicates present in the data. 2.2.1.1 Terms Related to Data Cleaning Data cleaning: data cleaning is the process of detecting, diagnosing, and editing damaged data. Data editing: data editing means changing the value of data which are incorrect. Data flow: data flow is defined as passing of recorded information through succeeding information carriers. Inliers: Inliers are data values falling inside the projected range. Outlier: outliers are data value falling outside the projected range. Robust estimation: evaluation of statistical parameters, using methods that are less responsive to the effect of outliers than more conventional methods are called robust method. 2.2.1.2 Definition: Data Cleaning Data cleaning is a process used to identify imprecise, incomplete, or irrational data and then improving the quality through correction of detected errors and omissions. This process may include format checks Completeness checks Reasonableness checks Limit checks Review of the data to identify outliers or other errors Assessment of data by subject area experts (e.g. taxonomic specialists). By this process suspected records are flagged, documented and checked subsequently. And finally these suspected records can be corrected. Sometimes validation checks also involve checking for compliance against applicable standards, rules, and conventions. The general framework for data cleaning given as: Define and determine error types; Search and identify error instances; Correct the errors; Document error instances and error types; and Modify data entry procedures to reduce future errors. Data cleaning process is referred by different people by a number of terms. It is a matter of preference what one uses. These terms include: Error Checking, Error Detection, Data Validation, Data Cleaning, Data Cleansing, Data Scrubbing and Error Correction. We use Data Cleaning to encompass three sub-processes, viz. Data checking and error detection; Data validation; and Error correction. A fourth improvement of the error prevention processes could perhaps be added. 2.2.1.3 Problems with Data Here we just note some key problems with data Missing data : This problem occur because of two main reasons Data are absent in source where it is expected to be present. Some times data is present are not available in appropriately form Detecting missing data is usually straightforward and simpler. Erroneous data: This problem occurs when a wrong value is recorded for a real world value. Detection of erroneous data can be quite difficult. (For instance the incorrect spelling of a name) Duplicated data : This problem occur because of two reasons Repeated entry of same real world entity with some different values Some times a real world entity may have different identifications. Repeat records are regular and frequently easy to detect. The different identification of the same real world entities can be a very hard problem to identify and solve. Heterogeneities: When data from different sources are brought together in one analysis problem heterogeneity may occur. Heterogeneity could be Structural heterogeneity arises when the data structures reflect different business usage Semantic heterogeneity arises when the meaning of data is different n each system that is being combined Heterogeneities are usually very difficult to resolve since because they usually involve a lot of contextual data that is not well defined as metadata. Information dependencies in the relationship between the different sets of attribute are commonly present. Wrong cleaning mechanisms can further damage the information in the data. Various analysis tools handle these problems in different ways. Commercial offerings are available that assist the cleaning process, but these are often problem specific. Uncertainty in information systems is a well-recognized hard problem. In following a very simple examples of missing and erroneous data is shown Extensive support for data cleaning must be provided by data warehouses. Data warehouses have high probability of â€Å"dirty data† since they load and continuously refresh huge amounts of data from a variety of sources. Since these data warehouses are used for strategic decision making therefore the correctness of their data is important to avoid wrong decisions. The ETL (Extraction, Transformation, and Loading) process for building a data warehouse is illustrated in following . Data transformations are related with schema or data translation and integration, and with filtering and aggregating data to be stored in the data warehouse. All data cleaning is classically performed in a separate data performance area prior to loading the transformed data into the warehouse. A large number of tools of varying functionality are available to support these tasks, but often a significant portion of the cleaning and transformation work has to be done manually or by low-level programs that are difficult to write and maintain. A data cleaning method should assure following: It should identify and eliminate all major errors and inconsistencies in an individual data sources and also when integrating multiple sources. Data cleaning should be supported by tools to bound manual examination and programming effort and it should be extensible so that can cover additional sources. It should be performed in association with schema related data transformations based on metadata. Data cleaning mapping functions should be specified in a declarative way and be reusable for other data sources. 2.2.1.4 Data Cleaning: Phases 1. Analysis: To identify errors and inconsistencies in the database there is a need of detailed analysis, which involves both manual inspection and automated analysis programs. This reveals where (most of) the problems are present. 2. Defining Transformation and Mapping Rules: After discovering the problems, this phase are related with defining the manner by which we are going to automate the solutions to clean the data. We will find various problems that translate to a list of activities as a result of analysis phase. Example: Remove all entries for J. Smith because they are duplicates of John Smith Find entries with `bule in colour field and change these to `blue. Find all records where the Phone number field does not match the pattern (NNNNN NNNNNN). Further steps for cleaning this data are then applied. Etc †¦ 3. Verification: In this phase we check and assess the transformation plans made in phase- 2. Without this step, we may end up making the data dirtier rather than cleaner. Since data transformation is the main step that actually changes the data itself so there is a need to be sure that the applied transformations will do it correctly. Therefore test and examine the transformation plans very carefully. Example: Let we have a very thick C++ book where it says strict in all the places where it should say struct 4. Transformation: Now if it is sure that cleaning will be done correctly, then apply the transformation verified in last step. For large database, this task is supported by a variety of tools Backflow of Cleaned Data: In a data mining the main objective is to convert and move clean data into target system. This asks for a requirement to purify legacy data. Cleansing can be a complicated process depending on the technique chosen and has to be designed carefully to achieve the objective of removal of dirty data. Some methods to accomplish the task of data cleansing of legacy system include: n Automated data cleansing n Manual data cleansing n The combined cleansing process 2.2.1.5 Missing Values Data cleaning addresses a variety of data quality problems, including noise and outliers, inconsistent data, duplicate data, and missing values. Missing values is one important problem to be addressed. Missing value problem occurs because many tuples may have no record for several attributes. For Example there is a customer sales database consisting of a whole bunch of records (lets say around 100,000) where some of the records have certain fields missing. Lets say customer income in sales data may be missing. Goal here is to find a way to predict what the missing data values should be (so that these can be filled) based on the existing data. Missing data may be due to following reasons Equipment malfunction Inconsistent with other recorded data and thus deleted Data not entered due to misunderstanding Certain data may not be considered important at the time of entry Not register history or changes of the data How to Handle Missing Values? Dealing with missing values is a regular question that has to do with the actual meaning of the data. There are various methods for handling missing entries 1. Ignore the data row. One solution of missing values is to just ignore the entire data row. This is generally done when the class label is not there (here we are assuming that the data mining goal is classification), or many attributes are missing from the row (not just one). But if the percentage of such rows is high we will definitely get a poor performance. 2. Use a global constant to fill in for missing values. We can fill in a global constant for missing values such as unknown, N/A or minus infinity. This is done because at times is just doesnt make sense to try and predict the missing value. For example if in customer sales database if, say, office address is missing for some, filling it in doesnt make much sense. This method is simple but is not full proof. 3. Use attribute mean. Let say if the average income of a a family is X you can use that value to replace missing income values in the customer sales database. 4. Use attribute mean for all samples belonging to the same class. Lets say you have a cars pricing DB that, among other things, classifies cars to Luxury and Low budget and youre dealing with missing values in the cost field. Replacing missing cost of a luxury car with the average cost of all luxury cars is probably more accurate then the value youd get if you factor in the low budget 5. Use data mining algorithm to predict the value. The value can be determined using regression, inference based tools using Bayesian formalism, decision trees, clustering algorithms etc. 2.2.1.6 Noisy Data Noise can be defined as a random error or variance in a measured variable. Due to randomness it is very difficult to follow a strategy for noise removal from the data. Real world data is not always faultless. It can suffer from corruption which may impact the interpretations of the data, models created from the data, and decisions made based on the data. Incorrect attribute values could be present because of following reasons Faulty data collection instruments Data entry problems Duplicate records Incomplete data: Inconsistent data Incorrect processing Data transmission problems Technology limitation. Inconsistency in naming convention Outliers How to handle Noisy Data? The methods for removing noise from data are as follows. 1. Binning: this approach first sort data and partition it into (equal-frequency) bins then one can smooth it using- Bin means, smooth using bin median, smooth using bin boundaries, etc. 2. Regression: in this method smoothing is done by fitting the data into regression functions. 3. Clustering: clustering detect and remove outliers from the data. 4. Combined computer and human inspection: in this approach computer detects suspicious values which are then checked by human experts (e.g., this approach deal with possible outliers).. Following methods are explained in detail as follows: Binning: Data preparation activity that converts continuous data to discrete data by replacing a value from a continuous range with a bin identifier, where each bin represents a range of values. For instance, age can be changed to bins such as 20 or under, 21-40, 41-65 and over 65. Binning methods smooth a sorted data set by consulting values around it. This is therefore called local smoothing. Let consider a binning example Binning Methods n Equal-width (distance) partitioning Divides the range into N intervals of equal size: uniform grid if A and B are the lowest and highest values of the attribute, the width of intervals will be: W = (B-A)/N. The most straightforward, but outliers may dominate presentation Skewed data is not handled well n Equal-depth (frequency) partitioning 1. It divides the range (values of a given attribute) into N intervals, each containing approximately same number of samples (elements) 2. Good data scaling 3. Managing categorical attributes can be tricky. n Smooth by bin means- Each bin value is replaced by the mean of values n Smooth by bin medians- Each bin value is replaced by the median of values n Smooth by bin boundaries Each bin value is replaced by the closest boundary value Example Let Sorted data for price (in dollars): 4, 8, 9, 15, 21, 21, 24, 25, 26, 28, 29, 34 n Partition into equal-frequency (equi-depth) bins: o Bin 1: 4, 8, 9, 15 o Bin 2: 21, 21, 24, 25 o Bin 3: 26, 28, 29, 34 n Smoothing by bin means: o Bin 1: 9, 9, 9, 9 ( for example mean of 4, 8, 9, 15 is 9) o Bin 2: 23, 23, 23, 23 o Bin 3: 29, 29, 29, 29 n Smoothing by bin boundaries: o Bin 1: 4, 4, 4, 15 o Bin 2: 21, 21, 25, 25 o Bin 3: 26, 26, 26, 34 Regression: Regression is a DM technique used to fit an equation to a dataset. The simplest form of regression is linear regression which uses the formula of a straight line (y = b+ wx) and determines the suitable values for b and w to predict the value of y based upon a given value of x. Sophisticated techniques, such as multiple regression, permit the use of more than one input variable and allow for the fitting of more complex models, such as a quadratic equation. Regression is further described in subsequent chapter while discussing predictions. Clustering: clustering is a method of grouping data into different groups , so that data in each group share similar trends and patterns. Clustering constitute a major class of data mining algorithms. These algorithms automatically partitions the data space into set of regions or cluster. The goal of the process is to find all set of similar examples in data, in some optimal fashion. Following shows three clusters. Values that fall outsid e the cluster are outliers. 4. Combined computer and human inspection: These methods find the suspicious values using the computer programs and then they are verified by human experts. By this process all outliers are checked. 2.2.1.7 Data cleaning as a process Data cleaning is the process of Detecting, Diagnosing, and Editing Data. Data cleaning is a three stage method involving repeated cycle of screening, diagnosing, and editing of suspected data abnormalities. Many data errors are detected by the way during study activities. However, it is more efficient to discover inconsistencies by actively searching for them in a planned manner. It is not always right away clear whether a data point is erroneous. Many times it requires careful examination. Likewise, missing values require additional check. Therefore, predefined rules for dealing with errors and true missing and extreme values are part of good practice. One can monitor for suspect features in survey questionnaires, databases, or analysis data. In small studies, with the examiner intimately involved at all stages, there may be small or no difference between a database and an analysis dataset. During as well as after treatment, the diagnostic and treatment phases of cleaning need insight into the sources and types of errors at all stages of the study. Data flow concept is therefore crucial in this respect. After measurement the research data go through repeated steps of- entering into information carriers, extracted, and transferred to other carriers, edited, selected, transformed, summarized, and presented. It is essential to understand that errors can occur at any stage of the data flow, including during data cleaning itself. Most of these problems are due to human error. Inaccuracy of a single data point and measurement may be tolerable, and associated to the inherent technological error of the measurement device. Therefore the process of data clenaning mus focus on those errors that are beyond small technical variations and that form a major shift within or beyond the population distribution. In turn, it must be based on understanding of technical errors and expected ranges of normal values. Some errors are worthy of higher priority, but which ones are most significant is highly study-specific. For instance in most medical epidemiological studies, errors that need to be cleaned, at all costs, include missing gender, gender misspecification, birth date or examination date errors, duplications or merging of records, and biologically impossible results. Another example is in nutrition studies, date errors lead to age errors, which in turn lead to errors in weight-for-age scoring and, further, to misclassification of subjects as under- or overweight. Errors of sex and date are particularly important because they contaminate derived variables. Prioritization is essential if the study is under time pressures or if resources for data cleaning are limited. 2.2.2 Data Integration This is a process of taking data from one or more sources and mapping it, field by field, onto a new data structure. Idea is to combine data from multiple sources into a coherent form. Various data mining projects requires data from multiple sources because n Data may be distributed over different databases or data warehouses. (for example an epidemiological study that needs information about hospital admissions and car accidents) n Sometimes data may be required from different geographic distributions, or there may be need for historical data. (e.g. integrate historical data into a new data warehouse) n There may be a necessity of enhancement of data with additional (external) data. (for improving data mining precision) 2.2.2.1 Data Integration Issues There are number of issues in data integrations. Consider two database tables. Imagine two database tables Database Table-1 Database Table-2 In integration of there two tables there are variety of issues involved such as 1. The same attribute may have different names (for example in above tables Name and Given Name are same attributes with different names) 2. An attribute may be derived from another (for example attribute Age is derived from attribute DOB) 3. Attributes might be redundant( For example attribute PID is redundant) 4. Values in attributes might be different (for example for PID 4791 values in second and third field are different in both the tables) 5. Duplicate records under different keys( there is a possibility of replication of same record with different key values) Therefore schema integration and object matching can be trickier. Question here is how equivalent entities from different sources are matched? This problem is known as entity identification problem. Conflicts have to be detected and resolved. Integration becomes easier if unique entity keys are available in all the data sets (or tables) to be linked. Metadata can help in schema integration (example of metadata for each attribute includes the name, meaning, data type and range of values permitted for the attribute) 2.2.2.1 Redundancy Redundancy is another important issue in data integration. Two given attribute (such as DOB and age for instance in give table) may be redundant if one is derived form the other attribute or set of attributes. Inconsistencies in attribute or dimension naming can lead to redundancies in the given data sets. Handling Redundant Data We can handle data redundancy problems by following ways n Use correlation analysis n Different coding / representation has to be considered (e.g. metric / imperial measures) n Careful (manual) integration of the data can reduce or prevent redundancies (and inconsistencies) n De-duplication (also called internal data linkage) o If no unique entity keys are available o Analysis of values in attributes to find duplicates n Process redundant and inconsistent data (easy if values are the same) o Delete one of the values o Average values (only for numerical attributes) o Take majority values (if more than 2 duplicates and some values are the same) Correlation analysis is explained in detail here. Correlation analysis (also called Pearsons product moment coefficient): some redundancies can be detected by using correlation analysis. Given two attributes, such analysis can measure how strong one attribute implies another. For numerical attribute we can compute correlation coefficient of two attributes A and B to evaluate the correlation between them. This is given by Where n n is the number of tuples, n and are the respective means of A and B n ÏÆ'A and ÏÆ'B are the respective standard deviation of A and B n ÃŽ £(AB) is the sum of the AB cross-product. a. If -1 b. If rA, B is equal to zero it indicates A and B are independent of each other and there is no correlation between them. c. If rA, B is less than zero then A and B are negatively correlated. , where if value of one attribute increases value of another attribute decreases. This means that one attribute discourages another attribute. It is important to note that correlation does not imply causality. That is, if A and B are correlated, this does not essentially mean that A causes B or that B causes A. for example in analyzing a demographic database, we may find that attribute representing number of accidents and the number of car theft in a region are correlated. This does not mean that one is related to another. Both may be related to third attribute, namely population. For discrete data, a correlation relation between two attributes, can be discovered by a χ ²(chi-square) test. Let A has c distinct values a1,a2,†¦Ã¢â‚¬ ¦ac and B has r different values namely b1,b2,†¦Ã¢â‚¬ ¦br The data tuple described by A and B are shown as contingency table, with c values of A (making up columns) and r values of B( making up rows). Each and every (Ai, Bj) cell in table has. X^2 = sum_{i=1}^{r} sum_{j=1}^{c} {(O_{i,j} E_{i,j})^2 over E_{i,j}} . Where n Oi, j is the observed frequency (i.e. actual count) of joint event (Ai, Bj) and n Ei, j is the expected frequency which can be computed as E_{i,j}=frac{sum_{k=1}^{c} O_{i,k} sum_{k=1}^{r} O_{k,j}}{N} , , Where n N is number of data tuple n Oi,k is number of tuples having value ai for A n Ok,j is number of tuples having value bj for B The larger the χ ² value, the more likely the variables are related. The cells that contribute the most to the χ ² value are those whose actual count is very different from the expected count Chi-Square Calculation: An Example Suppose a group of 1,500 people were surveyed. The gender of each person was noted. Each person has polled their preferred type of reading material as fiction or non-fiction. The observed frequency of each possible joint event is summarized in following table.( number in parenthesis are expected frequencies) . Calculate chi square. Play chess Not play chess Sum (row) Like science fiction 250(90) 200(360) 450 Not like science fiction 50(210) 1000(840) 1050 Sum(col.) 300 1200 1500 E11 = count (male)*count(fiction)/N = 300 * 450 / 1500 =90 and so on For this table the degree of freedom are (2-1)(2-1) =1 as table is 2X2. for 1 degree of freedom , the χ ² value needed to reject the hypothesis at the 0.001 significance level is 10.828 (taken from the table of upper percentage point of the χ ² distribution typically available in any statistic text book). Since the computed value is above this, we can reject the hypothesis that gender and preferred reading are independent and conclude that two attributes are strongly correlated for given group. Duplication must also be detected at the tuple level. The use of renormalized tables is also a source of redundancies. Redundancies may further lead to data inconsistencies (due to updating some but not others). 2.2.2.2 Detection and resolution of data value conflicts Another significant issue in data integration is the discovery and resolution of data value conflicts. For example, for the same entity, attribute values from different sources may differ. For example weight can be stored in metric unit in one source and British imperial unit in another source. For instance, for a hotel cha