《scikit-learn 快速入门》

《scikit-learn 快速入门》

scikit-learn 是 Python 最常用的传统机器学习库,提供分类、回归、聚类、降维、预处理等完整工具链。本文介绍安装、数据加载、模型训练与评估的基础流程。

1 安装与文档

pip install scikit-learn

2 基础流程:训练一个分类器

from sklearn.datasets import load_iris from sklearn.model_selection import train_test_split from sklearn.ensemble import RandomForestClassifier from sklearn.metrics import accuracy_score, classification_report # 1. 加载数据 iris = load_iris() X, y = iris.data, iris.target # 2. 划分训练集 / 测试集 X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, random_state=42 ) # 3. 训练模型(fit) clf = RandomForestClassifier(n_estimators=100, random_state=42) clf.fit(X_train, y_train) # 4. 预测与评估(predict / score) y_pred = clf.predict(X_test) print("accuracy:", accuracy_score(y_test, y_pred)) print(classification_report(y_test, y_pred)) # 训练集评分 print("train score:", clf.score(X_train, y_train))

3 数据预处理

from sklearn.preprocessing import StandardScaler, MinMaxScaler, LabelEncoder # 数值特征标准化(均值 0,方差 1)——很多模型(SVM、KNN)需要 scaler = StandardScaler() X_scaled = scaler.fit_transform(X) # 归一化到 [0, 1] scaler = MinMaxScaler() X_norm = scaler.fit_transform(X) # 标签编码(字符串 -> 整数) le = LabelEncoder() y_encoded = le.fit_transform(["cat", "dog", "cat"])

4 Pipeline 组合流程

用 Pipeline 把预处理与模型串起来,避免数据泄漏、代码更整洁:

from sklearn.pipeline import Pipeline pipe = Pipeline([ ("scaler", StandardScaler()), ("clf", RandomForestClassifier(n_estimators=100, random_state=42)), ]) pipe.fit(X_train, y_train) print("pipeline acc:", pipe.score(X_test, y_test))

5 交叉验证与超参调优

from sklearn.model_selection import cross_val_score, GridSearchCV # 5.1 交叉验证 scores = cross_val_score(clf, X, y, cv=5) # 5 折 print("cv scores:", scores, "mean:", scores.mean()) # 5.2 网格搜索找最佳超参 param_grid = { "n_estimators": [50, 100, 200], "max_depth": [None, 5, 10], } grid = GridSearchCV(RandomForestClassifier(random_state=42), param_grid, cv=5, scoring="accuracy", n_jobs=-1) grid.fit(X_train, y_train) print("best params:", grid.best_params_) print("best score:", grid.best_score_)

6 常用模型速查

任务 常用模型
分类 LogisticRegression、RandomForestClassifier、SVC、GradientBoostingClassifier、XGBoost(第三方)
回归 LinearRegression、Ridge、RandomForestRegressor
聚类 KMeans、DBSCAN、AgglomerativeClustering
降维 PCA、TruncatedSVD

7 常见问题

  • 特征尺度差异大:SVM、KNN、PCA 前务必先 StandardScaler。
  • 类别不平衡:用 class_weight="balanced" 或 imbalanced-learn 的 RandomUnderSampler。
  • 数据泄漏:fit_transform 只应在训练集上;测试集统一用训练集拟合好的 scaler 做 transform,Pipeline 可自动保证这一点。

参考文档

阅读 — · 全站 —
🎸 我的歌单 0 首