4 views

1 Answers

Data mining is the process of extracting and discovering patterns in large data sets involving methods at the intersection of machine learning, statistics, and database systems. Data mining is an interdisciplinary subfield of computer science and statistics with an overall goal of extracting information from a data set and transforming the information into a comprehensible structure for further use. Data mining is the analysis step of the "knowledge discovery in databases" process, or KDD. Aside from the raw analysis step, it also involves database and data management aspects, data pre-processing, model and inference considerations, interestingness metrics, complexity considerations, post-processing of discovered structures, visualization, and online updating.

The term "data mining" is a misnomer because the goal is the extraction of patterns and knowledge from large amounts of data, not the extraction of data itself. It also is a buzzword and is frequently applied to any form of large-scale data or information processing as well as any application of computer decision support system, including artificial intelligence and business intelligence. The book Data mining: Practical machine learning tools and techniques with Java was originally to be named Practical machine learning, and the term data mining was only added for marketing reasons. Often the more general terms data analysis and analytics—or, when referring to actual methods, artificial intelligence and machine learning—are more appropriate.

The actual data mining task is the semi-automatic or automatic analysis of large quantities of data to extract previously unknown, interesting patterns such as groups of data records , unusual records , and dependencies. This usually involves using database techniques such as spatial indices. These patterns can then be seen as a kind of summary of the input data, and may be used in further analysis or, for example, in machine learning and predictive analytics. For example, the data mining step might identify multiple groups in the data, which can then be used to obtain more accurate prediction results by a decision support system. Neither the data collection, data preparation, nor result interpretation and reporting is part of the data mining step, although they do belong to the overall KDD process as additional steps.

The difference between data analysis and data mining is that data analysis is used to test models and hypotheses on the dataset, e.g., analyzing the effectiveness of a marketing campaign, regardless of the amount of data. In contrast, data mining uses machine learning and statistical models to uncover clandestine or hidden patterns in a large volume of data.

4 views

Related Questions

What is Mining law?
1 Answers 4 Views
What is Salt mining?
1 Answers 4 Views
What is Uranium mining debate?
1 Answers 4 Views
What is Placer mining?
1 Answers 4 Views
What is Mining accident?
1 Answers 4 Views
What is Outburst (mining)?
1 Answers 4 Views
What is Hoist (mining)?
1 Answers 4 Views