Skills Plugins MCP Prompt Model 博客 我的中心
数据分析与咨询 #data #automation

data-cleaning-pipeline

Build robust processes for data cleaning, missing value imputation, outlier handling, and data transformation for data preprocessing, data quality, and data pipeline automation

DeepseekModel 官方收录技能 质量 优秀 · 78 v1.0.0

获取

https://deepseekmodel.com/api/download.php?id=aj-geddes-useful-ai-prompts-skills-data-cleaning-pipeline-skill-md&format=skill
下载 .skill 标准格式,含 system_prompt 与 model_config,导入任意 Agent 框架即可使用
.skill 文件中 system_prompt 字段的实际内容。
name Data Cleaning Pipeline description Build robust processes for data cleaning, missing value imputation, outlier handling, and data transformation for data preprocessing, data quality, and data pipeline automation Data Cleaning Pipeline Overview Data cleaning pipelines transform raw, messy data into clean, standardized formats suitable for analysis and modeling through systematic handling of missing values, outliers, and data quality issues. When to Use Preparing raw datasets for analysis or modeling Handling missing values and data quality issues Removing duplicates and standardizing formats Detecting and treating outliers Building automated data preprocessing workflows Ensuring data integrity and consistency Core Components Missing Value Handling : Imputation and removal strategies Outlier Detection & Treatment : Identifying and handling anomalies Data Type Standardization : Ensuring correct data types Duplicate Removal : Identifying and removing duplicates Normalization & Scaling : Standardizing value ranges Text Cleaning : Handling text data Cleaning Strategies Deletion : Removing rows or columns Imputation : Filling with mean, median, or predictive models Transformation : Converting between formats Validation : Ensuring data integrity rules Implementation with Python import pandas as pd import numpy as np from sklearn.preprocessing import StandardScaler, MinMaxScaler from sklearn.impute import SimpleImputer, KNNImputer # Load raw data df = pd.read_csv( 'raw_data.csv' ) # Step 1: Identify and handle missing values print ( "Missing values:\n" , df.isnull(). sum ()) # Strategy 1: Delete rows with critical missing values df = df.dropna(subset=[ 'customer_id' , 'transaction_date' ]) # Strategy 2: Impute numerical columns with median imputer = SimpleImputer(strategy= 'median' ) df[ 'age' ] = imputer.fit_transform(df[[ 'age' ]]) # Strategy 3: Use KNN imputation for related features knn_imputer = KNNImputer(n_neighbors= 5 ) numeric_cols = df.select_dtypes(include=[np.number]).columns df[numeric_cols] = knn_imputer.fit_transform(df[numeric_cols]) # Strategy 4: Fill categorical with mode df[ 'category' ] = df[ 'category' ].fillna(df[ 'category' ].mode()[ 0 ]) # Step 2: Handle duplicates print ( f"Duplicate rows: {df.duplicated(). sum ()} " ) df = df.drop_duplicates() # Duplicate on specific columns df = df.drop_duplicates(subset=[ 'customer_id' , 'transaction_date' ]) # Step 3: Outlier detection and handling Q1 = df[ 'amount' ].quantile( 0.25 ) Q3 = df[ 'amount' ].quantile( 0.75 ) IQR = Q3 - Q1 lower_bound = Q1 - 1.5 * IQR upper_bound = Q3 + 1.5 * IQR # Remove outliers df = df[(df[ 'amount' ] >= lower_bound) & (df[ 'amount' ] <= upper_bound)] # Alternative: Cap outliers df[ 'amount' ] = df[ 'amount' ].clip(lower=lower_bound, upper=upper_bound) # Step 4: Data type standardization df[ 'transaction_date' ] = pd.to_datetime(df[ 'transaction_date' ]) df[ 'customer_id' ] = df[ 'customer_id' ].astype( 'int64' ) df[ 'amount' ] = pd.to_numeric(df[ 'amount' ], errors= 'coerce' ) # Step 5: Text cleaning df[ 'name' ] = df[ 'name' ]. str .strip(). str .lower() df[ 'name' ] = df[ 'name' ]. str .replace( '[^a-z0-9\s]' , '' , regex= True ) # Step 6: Normalization and scaling scaler = StandardScaler() df[[ 'age' , 'income' ]] = scaler.fit_transform(df[[ 'age' , 'income' ]]) # MinMax scaling for bounded range [0, 1] minmax_scaler = MinMaxScaler() df[[ 'score' ]] = minmax_scaler.fit_transform(df[[ 'score' ]]) # Step 7: Create data quality report def create_quality_report ( df_original, df_cleaned ): report = { 'Original rows' : len (df_original), 'Cleaned rows' : len (df_cleaned), 'Rows removed' : len (df_original) - len (df_cleaned), 'Removal percentage' : (( len (df_original) - len (df_cleaned)) / len (df_original) * 100 ), 'Original missing' : df_original.isnull(). sum (). sum (), 'Cleaned missing' : df_cleaned.isnull(). sum (). sum (), } return pd.DataFrame(report, index=[ 0 ]) quality = create_quality_report(df, df) print (quality) # Step 8: Validation checks assert df[ 'age' ].isnull(). sum () == 0 , "Age has missing values" assert df[ 'transaction_date' ].dtype == 'datetime64[ns]' , "Date not datetime" assert (df[ 'amount' ] >= 0 ). all (), "Negative amounts detected" print ( "Data cleaning pipeline completed successfully!" ) Pipeline Architecture class DataCleaningPipeline : def __init__ ( self ): self .cleaner_steps = [] def add_step ( self, func, description ): self .cleaner_steps.append((func, description)) return self def execute ( self, df ): for func, desc in self .cleaner_steps: print ( f"Executing: {desc} " ) df = func(df) return df # Usage pipeline = DataCleaningPipeline() pipeline.add_step( lambda df: df.dropna(subset=[ 'customer_id' ]), "Remove rows with missing customer_id" ).add_step( lambda df: df.drop_duplicates(), "Remove duplicate rows" ).add_step( lambda df: df[(df[ 'amount' ] > 0 ) & (df[ 'amount' ] < 100000 )], "Filter invalid amount ranges" ) df_clean = pipeline.execute(df) Advanced Cleaning Techniques # Step 9: Feature-specific cleaning df[ 'phone' ] = df[ 'phone' ]. str .replace( r'\D' , '' , regex= True ) # Remove non-digits # Step 10: Datetime handling df[ 'created_date' ] = pd.to_datetime(df[ 'created_date' ], errors= 'coerce' ) df[ 'days_since_creation' ] = (pd.Timestamp.now() - df[ 'created_date' ]).dt.days # Step 11: Categorical standardization df[ 'status' ] = df[ 'status' ]. str .lower(). str .strip() df[ 'status' ] = df[ 'status' ].replace({ 'active' : 'active' , 'inactive' : 'inactive' , 'pending' : 'pending' , }) # Step 12: Numeric constraint checking df[ 'age' ] = df[ 'age' ].where((df[ 'age' ] >= 0 ) & (df[ 'age' ] <= 150 ), np.nan) df[ 'percentage' ] = df[ 'percentage' ].where((df[ 'percentage' ] >= 0 ) & (df[ 'percentage' ] <= 100 ), np.nan) # Step 13: Create data quality score quality_score = { 'Missing %' : (df.isnull(). sum () / len (df) * 100 ).mean(), 'Duplicates %' : (df.duplicated(). sum () / len (df) * 100 ), 'Complete Features' : (df.notna(). sum () / len (df)).mean() * 100 , } # Step 14: Generate cleaning report cleaning_report = f""" DATA CLEANING REPORT ==================== Rows removed: { len (df) - len (df_clean)} Columns: { len (df_clean.columns)} Remaining rows: { len (df_clean)} Completeness: {(df_clean.notna(). sum (). sum () / ( len (df_clean) * len (df_clean.columns)) * 100 ): .1 f} % """ print (cleaning_report) Key Decisions How to handle missing values (delete vs impute)? Which outliers are legitimate business cases? What are acceptable value ranges? Which duplicates are true duplicates? How to standardize categorical values? Validation Steps Check for data type consistency Verify value ranges are reasonable Confirm no unintended data loss Document all transformations applied Create audit trail of changes Deliverables Cleaned dataset with quality metrics Data cleaning log documenting all steps Validation report confirming data integrity Before/after comparison statistics Cleaning code and pipeline documentation
Agent 识别该技能的关键词,点击任意一个即可复制。

该技能未提供触发词。

下载的 .skill 包内含以下字段。
字段 说明
format格式标识(skill/v1)
skill_id技能唯一 ID
name技能名称
version版本号
description技能描述
category所属分类(数组)
trigger_words触发词列表
tags标签列表
source来源标识
source_url来源链接(本页地址)
exported_at导出时间(每次下载生成)
system_prompt系统提示词正文
model_config模型参数:provider / model / temperature / max_tokens / top_p
examples示例
install_guide各平台导入说明(Coze / Dify / Claude / 自定义框架)
同一份技能可按不同平台格式导出。
.skill 标准格式,含 system_prompt 与 model_config,导入任意 Agent 框架即可使用 下载
.skillpro 增强格式,额外含脚本 / 工具 / 依赖 / 钩子占位 下载
.json 纯 JSON 导出,只含 system_prompt 与模型参数 下载
Coze 带 frontmatter 的 Markdown,Coze 平台导入用 下载
Dify Dify DSL,创建应用后直接导入 下载

每日精选 Skill 推荐,免费送到你邮箱

输入邮箱,每天接收一个精选 AI Agent 技能推荐。完全免费,持续更新。

验证码 --

提交后我们会发送一封确认邮件,点击邮件里的链接才会开始收信。

完全免费,取消任意时间。我们不会发送垃圾邮件。