{"cells":[{"metadata":{},"cell_type":"markdown","source":"# Mean Baseline"},{"metadata":{},"cell_type":"markdown","source":"fork from https://www.kaggle.com/paulorzp/mean-baseline  \nthe difference is that I calculate mean value of `rv_lv_ratio_gte_1, rv_lv_ratio_lt_1, leftsided_pe, rightsided_pe, central_pe, chronic_pe, acute_and_chronic_pe` for positive exam only.\n\nI modified so because \n`rv_lv_ratio_gte_1, rv_lv_ratio_lt_1, leftsided_pe, rightsided_pe, central_pe, chronic_pe, acute_and_chronic_pe` in train.csv is valid only when negative_exam_for_pe==0&&indeterminate==0 (I found above columns is all zero when negative_exam_for_pe==1||indeterminate==1)\n\n**but above modification make the score worse**  \noriginal (v3): 0.555  \nfixed(v2): 0.657\n\nI recheck the evaluation description and found that `rv_lv_ratio_gte_1, rv_lv_ratio_lt_1, leftsided_pe, rightsided_pe, central_pe, chronic_pe, acute_and_chronic_pe` is ALSO EVALUATED FOR ALL EXAMS INCLUDING NON-POSITIVE EXAMS. So this result is reasonable.\n\nThough it may be still ok to eval `rv_lv_ratio_gte_1, rv_lv_ratio_lt_1` for all exam, \nIt may be better to eval `leftsided_pe, rightsided_pe, central_pe, chronic_pe, acute_and_chronic_pe` for positive exam only because these columns are only valid for positive exams. Also, it is consistent with the fact that pe_present_on_image is evaluated for positive exam only.\n\nIt seems that scores for `leftsided_pe, rightsided_pe, central_pe, chronic_pe, acute_and_chronic_pe` is affected by exam level pos/neg prediction accuracy."},{"metadata":{"trusted":true},"cell_type":"code","source":"USE_POS_MEAN = False","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import pandas as pd","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"PATH = \"../input/rsna-str-pulmonary-embolism-detection/\"\ntrain = pd.read_csv(PATH + \"train.csv\")\nsub = pd.read_csv(PATH + \"sample_submission.csv\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"feats = list(train.columns[3:5])+list(train.columns[8:12])+list(train.columns[13:17])\nfeats","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"means = train[feats].mean().to_dict()\nmeans","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"means_pos_row = train.query(\"negative_exam_for_pe==0 and indeterminate==0\")[feats].mean().to_dict()\nmeans_pos_row","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# below values not used\n# columns exept negative_exam_for_pe and indeterminate are all zero\nmeans_not_pos_row = train.query(\"negative_exam_for_pe==1 or indeterminate==1\")[feats].mean().to_dict()\nmeans_not_pos_row","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"POS_ONLY_ROWS = [\"rv_lv_ratio_gte_1\", \"rv_lv_ratio_lt_1\", \"leftsided_pe\", \"rightsided_pe\", \"central_pe\", \"chronic_pe\", \"acute_and_chronic_pe\"]\n\n# pe_present_on_image prediction for sample_submission's id_only rows\nsub['label'] = means['pe_present_on_image']\n\nfor feat in means.keys():\n    sub.loc[sub.id.str.contains(feat, regex=False), 'label'] = means[feat]\n\nif USE_POS_MEAN:\n    for feat in POS_ONLY_ROWS:\n        sub.loc[sub.id.str.contains(feat, regex=False), 'label'] = means_pos_row[feat]\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sub.to_csv('submission.csv', index = False, float_format='%.4g')","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}