{
  "id": 447501,
  "title": "𝐄𝐱𝐭𝐫𝐚𝐯𝐚𝐬𝐚𝐭𝐢𝐨𝐧 𝐬𝐞𝐫𝐢𝐞𝐬 𝐜𝐥𝐚𝐬𝐬𝐢𝐟𝐢𝐜𝐚𝐭𝐢𝐨𝐧 𝐨𝐧 𝐘𝐎𝐋𝐎 𝐜𝐨𝐧𝐟𝐢𝐝𝐞𝐧𝐜𝐞 𝐚𝐩𝐩𝐫𝐨𝐚𝐜𝐡.",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/447501",
  "author_name": "Georgii Aparin",
  "post_date": "2023-10-16T07:50:15.950000",
  "votes": 14,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I made a rather interesting approach to classify extravasation and want to share it with you.</p>\n<p>My idea would have been impossible to implement without the bounding box <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/441402\" target=\"_blank\">dataset</a> from <a href=\")(https://www.kaggle.com/vaillant\" target=\"_blank\">Ian Pan</a> . Thanks a lot for his work.</p>\n<p>Using this dataset, I trained YOLO detection and collected my “time” series dataset. The idea was to collect confidence and area of ​​the bounding boxes. Walking through the sorted scans of the axial plane, I collected model predictions into my dataset.</p>\n<h1>𝐓𝐡𝐢𝐬 𝐢𝐬 𝐰𝐡𝐚𝐭 𝐭𝐡𝐞 “𝐭𝐢𝐦𝐞” 𝐬𝐞𝐫𝐢𝐞𝐬 𝐥𝐨𝐨𝐤𝐞𝐝 𝐥𝐢𝐤𝐞:</h1>\n<h2>𝐄𝐱𝐭𝐫𝐚𝐯𝐚𝐬𝐚𝐭𝐢𝐨𝐧</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11837581%2Fd1aff4084ad3b318e2a6350a4869fca3%2FScreenshot%202023-10-16%20at%2009.55.53.png?generation=1697440072235399&amp;alt=media\" alt=\"\"></p>\n<h2>𝐍𝐨 𝐞𝐱𝐭𝐫𝐚𝐯𝐚𝐬𝐚𝐭𝐢𝐨𝐧</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11837581%2F6235a05af397f1b3e2e0e42c01c9585e%2FScreenshot%202023-10-16%20at%2010.08.53.png?generation=1697440146957921&amp;alt=media\" alt=\"\"></p>\n<h1>𝐃𝐚𝐭𝐚 𝐩𝐫𝐞𝐩𝐚𝐫𝐢𝐧𝐠:</h1>\n<p>When assembling the dataset, I also experimented with TTA, but as practice has shown, this did not bring a big increase in quality, but took 4 times more time for inference.</p>\n<h1>𝐅𝐞𝐚𝐭𝐮𝐫𝐞 𝐞𝐧𝐠𝐢𝐧𝐞𝐞𝐫𝐢𝐧𝐠:</h1>\n<p>After that, I started generating features for these series. After many attempts, I came to the conclusion that the simplest features, such as std, mean, median, etc. There were already enough of them for the optimal metric. I couldn’t separate the classes more clearly.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11837581%2Fc33d7d159235bb1375e68bf64315da5d%2FScreenshot%202023-10-16%20at%2010.11.35.png?generation=1697440490402173&amp;alt=media\" alt=\"\"></p>\n<h1>𝐂𝐥𝐚𝐬𝐬𝐢𝐟𝐢𝐜𝐚𝐭𝐢𝐨𝐧 𝐦𝐨𝐝𝐞𝐥𝐬:</h1>\n<p>Using these features, I experimented with various models, but settled on the logistic regression ensemble and <a href=\"https://imbalanced-learn.org/stable/references/generated/imblearn.ensemble.BalancedRandomForestClassifier.html\" target=\"_blank\">lmblearn Balanced Random Forest Classifier</a>.</p>\n<p>I made stratified cross-validation.</p>\n<pre><code>Score = Metric(label = 5)\nval_scores = []\n\n i  range(N_SPLITS):\n    X_train = df[feature_cols][df.fold != i]\n    y_train = df[df.fold != i].label\n    X_val = df[feature_cols][df.fold == i]\n    y_val = df[df.fold == i].label\n\n    LR = LogisticRegression(=21,\n                            =,\n                            class_weight={0 : 1, 1 : 6},\n                             =,\n                            =0.9)\n\n    BRF = BalancedRandomForestClassifier(=100,\n                                         =,\n                                         =None,\n                                         =2,\n                                         =1,\n                                         =0.,\n                                         =,\n                                         =None,\n                                         =0.,\n                                         =,\n                                         =,\n                                         =,\n                                         =,\n                                         =21,\n                                         =0,\n                                         =,\n                                         class_weight={0 : 1, 1 : 5},\n                                         =0.,\n                                         =None\n                                        )\n\n    fit_LR = LR.fit(X_train, y_train)\n    fit_BRF = BRF.fit(X_train, y_train)\n\n    pred = np.array(0.5 * fit_BRF.predict(X_val) + 0.5 * fit_LR.predict(X_val), dtype = np.uint8)\n    f1 = f1_score(y_val, pred)\n\n     = np.array(y_val)\n    pred_LR = np.array(fit_LR.predict_proba(X_val))\n    pred_BRF = np.array(fit_BRF.predict_proba(X_val))\n    pred = 0.5 * pred_BRF + 0.5 * pred_LR\n    val_score = Score.get_score(, pred)\n    val_scores.append(val_score)\n\n    (f)\n    (f)\n    (f)\n    ()\n\n(f)\n</code></pre>\n<h1>𝐀𝐧𝐝 𝐠𝐨𝐭 𝐭𝐡𝐞 𝐟𝐨𝐥𝐥𝐨𝐰𝐢𝐧𝐠 𝐫𝐞𝐬𝐮𝐥𝐭𝐬:</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11837581%2F4c45ae6d9091e8491ed5ad6553e1a396%2FScreenshot%202023-10-16%20at%2010.44.42.png?generation=1697442310307315&amp;alt=media\" alt=\"\"></p>\n<p>This approach showed 0.02 better logloss than the best statistical approach.<br>\nIt seems that the idea can be improved, for example by collecting better data or generating more suitable features.<br>\nThank you for your attention, I look forward to your criticism and suggestions.<br>\nI'm waiting for your questions.</p>\n<p>&lt;3</p>",
  "messages": [
    {
      "id": 2484072,
      "postDate": "2023-10-16T07:50:15.950Z",
      "content": "<p>I made a rather interesting approach to classify extravasation and want to share it with you.</p>\n<p>My idea would have been impossible to implement without the bounding box <a href=\"https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/441402\" target=\"_blank\">dataset</a> from <a href=\")(https://www.kaggle.com/vaillant\" target=\"_blank\">Ian Pan</a> . Thanks a lot for his work.</p>\n<p>Using this dataset, I trained YOLO detection and collected my “time” series dataset. The idea was to collect confidence and area of ​​the bounding boxes. Walking through the sorted scans of the axial plane, I collected model predictions into my dataset.</p>\n<h1>𝐓𝐡𝐢𝐬 𝐢𝐬 𝐰𝐡𝐚𝐭 𝐭𝐡𝐞 “𝐭𝐢𝐦𝐞” 𝐬𝐞𝐫𝐢𝐞𝐬 𝐥𝐨𝐨𝐤𝐞𝐝 𝐥𝐢𝐤𝐞:</h1>\n<h2>𝐄𝐱𝐭𝐫𝐚𝐯𝐚𝐬𝐚𝐭𝐢𝐨𝐧</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11837581%2Fd1aff4084ad3b318e2a6350a4869fca3%2FScreenshot%202023-10-16%20at%2009.55.53.png?generation=1697440072235399&amp;alt=media\" alt=\"\"></p>\n<h2>𝐍𝐨 𝐞𝐱𝐭𝐫𝐚𝐯𝐚𝐬𝐚𝐭𝐢𝐨𝐧</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11837581%2F6235a05af397f1b3e2e0e42c01c9585e%2FScreenshot%202023-10-16%20at%2010.08.53.png?generation=1697440146957921&amp;alt=media\" alt=\"\"></p>\n<h1>𝐃𝐚𝐭𝐚 𝐩𝐫𝐞𝐩𝐚𝐫𝐢𝐧𝐠:</h1>\n<p>When assembling the dataset, I also experimented with TTA, but as practice has shown, this did not bring a big increase in quality, but took 4 times more time for inference.</p>\n<h1>𝐅𝐞𝐚𝐭𝐮𝐫𝐞 𝐞𝐧𝐠𝐢𝐧𝐞𝐞𝐫𝐢𝐧𝐠:</h1>\n<p>After that, I started generating features for these series. After many attempts, I came to the conclusion that the simplest features, such as std, mean, median, etc. There were already enough of them for the optimal metric. I couldn’t separate the classes more clearly.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11837581%2Fc33d7d159235bb1375e68bf64315da5d%2FScreenshot%202023-10-16%20at%2010.11.35.png?generation=1697440490402173&amp;alt=media\" alt=\"\"></p>\n<h1>𝐂𝐥𝐚𝐬𝐬𝐢𝐟𝐢𝐜𝐚𝐭𝐢𝐨𝐧 𝐦𝐨𝐝𝐞𝐥𝐬:</h1>\n<p>Using these features, I experimented with various models, but settled on the logistic regression ensemble and <a href=\"https://imbalanced-learn.org/stable/references/generated/imblearn.ensemble.BalancedRandomForestClassifier.html\" target=\"_blank\">lmblearn Balanced Random Forest Classifier</a>.</p>\n<p>I made stratified cross-validation.</p>\n<pre><code>Score = Metric(label = 5)\nval_scores = []\n\n i  range(N_SPLITS):\n    X_train = df[feature_cols][df.fold != i]\n    y_train = df[df.fold != i].label\n    X_val = df[feature_cols][df.fold == i]\n    y_val = df[df.fold == i].label\n\n    LR = LogisticRegression(=21,\n                            =,\n                            class_weight={0 : 1, 1 : 6},\n                             =,\n                            =0.9)\n\n    BRF = BalancedRandomForestClassifier(=100,\n                                         =,\n                                         =None,\n                                         =2,\n                                         =1,\n                                         =0.,\n                                         =,\n                                         =None,\n                                         =0.,\n                                         =,\n                                         =,\n                                         =,\n                                         =,\n                                         =21,\n                                         =0,\n                                         =,\n                                         class_weight={0 : 1, 1 : 5},\n                                         =0.,\n                                         =None\n                                        )\n\n    fit_LR = LR.fit(X_train, y_train)\n    fit_BRF = BRF.fit(X_train, y_train)\n\n    pred = np.array(0.5 * fit_BRF.predict(X_val) + 0.5 * fit_LR.predict(X_val), dtype = np.uint8)\n    f1 = f1_score(y_val, pred)\n\n     = np.array(y_val)\n    pred_LR = np.array(fit_LR.predict_proba(X_val))\n    pred_BRF = np.array(fit_BRF.predict_proba(X_val))\n    pred = 0.5 * pred_BRF + 0.5 * pred_LR\n    val_score = Score.get_score(, pred)\n    val_scores.append(val_score)\n\n    (f)\n    (f)\n    (f)\n    ()\n\n(f)\n</code></pre>\n<h1>𝐀𝐧𝐝 𝐠𝐨𝐭 𝐭𝐡𝐞 𝐟𝐨𝐥𝐥𝐨𝐰𝐢𝐧𝐠 𝐫𝐞𝐬𝐮𝐥𝐭𝐬:</h1>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11837581%2F4c45ae6d9091e8491ed5ad6553e1a396%2FScreenshot%202023-10-16%20at%2010.44.42.png?generation=1697442310307315&amp;alt=media\" alt=\"\"></p>\n<p>This approach showed 0.02 better logloss than the best statistical approach.<br>\nIt seems that the idea can be improved, for example by collecting better data or generating more suitable features.<br>\nThank you for your attention, I look forward to your criticism and suggestions.<br>\nI'm waiting for your questions.</p>\n<p>&lt;3</p>",
      "rawMarkdown": "I made a rather interesting approach to classify extravasation and want to share it with you.\n\nMy idea would have been impossible to implement without the bounding box [dataset](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/441402) from [Ian Pan]()(https://www.kaggle.com/vaillant) . Thanks a lot for his work.\n\nUsing this dataset, I trained YOLO detection and collected my “time” series dataset. The idea was to collect confidence and area of ​​the bounding boxes. Walking through the sorted scans of the axial plane, I collected model predictions into my dataset.\n\n# 𝐓𝐡𝐢𝐬 𝐢𝐬 𝐰𝐡𝐚𝐭 𝐭𝐡𝐞 “𝐭𝐢𝐦𝐞” 𝐬𝐞𝐫𝐢𝐞𝐬 𝐥𝐨𝐨𝐤𝐞𝐝 𝐥𝐢𝐤𝐞:\n\n## 𝐄𝐱𝐭𝐫𝐚𝐯𝐚𝐬𝐚𝐭𝐢𝐨𝐧\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11837581%2Fd1aff4084ad3b318e2a6350a4869fca3%2FScreenshot%202023-10-16%20at%2009.55.53.png?generation=1697440072235399&alt=media)\n\n\n## 𝐍𝐨 𝐞𝐱𝐭𝐫𝐚𝐯𝐚𝐬𝐚𝐭𝐢𝐨𝐧\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11837581%2F6235a05af397f1b3e2e0e42c01c9585e%2FScreenshot%202023-10-16%20at%2010.08.53.png?generation=1697440146957921&alt=media)\n\n# 𝐃𝐚𝐭𝐚 𝐩𝐫𝐞𝐩𝐚𝐫𝐢𝐧𝐠:\nWhen assembling the dataset, I also experimented with TTA, but as practice has shown, this did not bring a big increase in quality, but took 4 times more time for inference.\n\n# 𝐅𝐞𝐚𝐭𝐮𝐫𝐞 𝐞𝐧𝐠𝐢𝐧𝐞𝐞𝐫𝐢𝐧𝐠:\nAfter that, I started generating features for these series. After many attempts, I came to the conclusion that the simplest features, such as std, mean, median, etc. There were already enough of them for the optimal metric. I couldn’t separate the classes more clearly.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11837581%2Fc33d7d159235bb1375e68bf64315da5d%2FScreenshot%202023-10-16%20at%2010.11.35.png?generation=1697440490402173&alt=media)\n\n# 𝐂𝐥𝐚𝐬𝐬𝐢𝐟𝐢𝐜𝐚𝐭𝐢𝐨𝐧 𝐦𝐨𝐝𝐞𝐥𝐬:\nUsing these features, I experimented with various models, but settled on the logistic regression ensemble and [lmblearn Balanced Random Forest Classifier](https://imbalanced-learn.org/stable/references/generated/imblearn.ensemble.BalancedRandomForestClassifier.html).\n\nI made stratified cross-validation.\n```\nScore = Metric(label = 5)\nval_scores = []\n\nfor i in range(N_SPLITS):\n    X_train = df[feature_cols][df.fold != i]\n    y_train = df[df.fold != i].label\n    X_val = df[feature_cols][df.fold == i]\n    y_val = df[df.fold == i].label\n    \n    LR = LogisticRegression(random_state=21,\n                            penalty='elasticnet',\n                            class_weight={0 : 1, 1 : 6},\n                             solver=\"saga\",\n                            l1_ratio=0.9)\n\n    BRF = BalancedRandomForestClassifier(n_estimators=100,\n                                         criterion=\"gini\",\n                                         max_depth=None,\n                                         min_samples_split=2,\n                                         min_samples_leaf=1,\n                                         min_weight_fraction_leaf=0.,\n                                         max_features='sqrt',\n                                         max_leaf_nodes=None,\n                                         min_impurity_decrease=0.,\n                                         bootstrap=True,\n                                         oob_score=False,\n                                         sampling_strategy=\"auto\",\n                                         replacement=False,\n                                         random_state=21,\n                                         verbose=0,\n                                         warm_start=False,\n                                         class_weight={0 : 1, 1 : 5},\n                                         ccp_alpha=0.,\n                                         max_samples=None\n                                        )\n\n    fit_LR = LR.fit(X_train, y_train)\n    fit_BRF = BRF.fit(X_train, y_train)\n    \n    pred = np.array(0.5 * fit_BRF.predict(X_val) + 0.5 * fit_LR.predict(X_val), dtype = np.uint8)\n    f1 = f1_score(y_val, pred)\n    \n    true = np.array(y_val)\n    pred_LR = np.array(fit_LR.predict_proba(X_val))\n    pred_BRF = np.array(fit_BRF.predict_proba(X_val))\n    pred = 0.5 * pred_BRF + 0.5 * pred_LR\n    val_score = Score.get_score(true, pred)\n    val_scores.append(val_score)\n    \n    print(f'fold: {i}')\n    print(f'f1 val score: {f1}')\n    print(f'w_logloss val score: {val_score}')\n    print()\n    \nprint(f\"mean val score: {np.mean(val_scores)} +- {2 * np.std(val_scores)}\")\n```\n\n# 𝐀𝐧𝐝 𝐠𝐨𝐭 𝐭𝐡𝐞 𝐟𝐨𝐥𝐥𝐨𝐰𝐢𝐧𝐠 𝐫𝐞𝐬𝐮𝐥𝐭𝐬:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11837581%2F4c45ae6d9091e8491ed5ad6553e1a396%2FScreenshot%202023-10-16%20at%2010.44.42.png?generation=1697442310307315&alt=media)\n\n\n\nThis approach showed 0.02 better logloss than the best statistical approach.\nIt seems that the idea can be improved, for example by collecting better data or generating more suitable features.\nThank you for your attention, I look forward to your criticism and suggestions.\nI'm waiting for your questions.\n\n<3\n\n\n",
      "votes": 13
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2484072": "I made a rather interesting approach to classify extravasation and want to share it with you.\n\nMy idea would have been impossible to implement without the bounding box [dataset](https://www.kaggle.com/competitions/rsna-2023-abdominal-trauma-detection/discussion/441402) from [Ian Pan]()(https://www.kaggle.com/vaillant) . Thanks a lot for his work.\n\nUsing this dataset, I trained YOLO detection and collected my “time” series dataset. The idea was to collect confidence and area of ​​the bounding boxes. Walking through the sorted scans of the axial plane, I collected model predictions into my dataset.\n\n# 𝐓𝐡𝐢𝐬 𝐢𝐬 𝐰𝐡𝐚𝐭 𝐭𝐡𝐞 “𝐭𝐢𝐦𝐞” 𝐬𝐞𝐫𝐢𝐞𝐬 𝐥𝐨𝐨𝐤𝐞𝐝 𝐥𝐢𝐤𝐞:\n\n## 𝐄𝐱𝐭𝐫𝐚𝐯𝐚𝐬𝐚𝐭𝐢𝐨𝐧\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11837581%2Fd1aff4084ad3b318e2a6350a4869fca3%2FScreenshot%202023-10-16%20at%2009.55.53.png?generation=1697440072235399&alt=media)\n\n\n## 𝐍𝐨 𝐞𝐱𝐭𝐫𝐚𝐯𝐚𝐬𝐚𝐭𝐢𝐨𝐧\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11837581%2F6235a05af397f1b3e2e0e42c01c9585e%2FScreenshot%202023-10-16%20at%2010.08.53.png?generation=1697440146957921&alt=media)\n\n# 𝐃𝐚𝐭𝐚 𝐩𝐫𝐞𝐩𝐚𝐫𝐢𝐧𝐠:\nWhen assembling the dataset, I also experimented with TTA, but as practice has shown, this did not bring a big increase in quality, but took 4 times more time for inference.\n\n# 𝐅𝐞𝐚𝐭𝐮𝐫𝐞 𝐞𝐧𝐠𝐢𝐧𝐞𝐞𝐫𝐢𝐧𝐠:\nAfter that, I started generating features for these series. After many attempts, I came to the conclusion that the simplest features, such as std, mean, median, etc. There were already enough of them for the optimal metric. I couldn’t separate the classes more clearly.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11837581%2Fc33d7d159235bb1375e68bf64315da5d%2FScreenshot%202023-10-16%20at%2010.11.35.png?generation=1697440490402173&alt=media)\n\n# 𝐂𝐥𝐚𝐬𝐬𝐢𝐟𝐢𝐜𝐚𝐭𝐢𝐨𝐧 𝐦𝐨𝐝𝐞𝐥𝐬:\nUsing these features, I experimented with various models, but settled on the logistic regression ensemble and [lmblearn Balanced Random Forest Classifier](https://imbalanced-learn.org/stable/references/generated/imblearn.ensemble.BalancedRandomForestClassifier.html).\n\nI made stratified cross-validation.\n```\nScore = Metric(label = 5)\nval_scores = []\n\nfor i in range(N_SPLITS):\n    X_train = df[feature_cols][df.fold != i]\n    y_train = df[df.fold != i].label\n    X_val = df[feature_cols][df.fold == i]\n    y_val = df[df.fold == i].label\n    \n    LR = LogisticRegression(random_state=21,\n                            penalty='elasticnet',\n                            class_weight={0 : 1, 1 : 6},\n                             solver=\"saga\",\n                            l1_ratio=0.9)\n\n    BRF = BalancedRandomForestClassifier(n_estimators=100,\n                                         criterion=\"gini\",\n                                         max_depth=None,\n                                         min_samples_split=2,\n                                         min_samples_leaf=1,\n                                         min_weight_fraction_leaf=0.,\n                                         max_features='sqrt',\n                                         max_leaf_nodes=None,\n                                         min_impurity_decrease=0.,\n                                         bootstrap=True,\n                                         oob_score=False,\n                                         sampling_strategy=\"auto\",\n                                         replacement=False,\n                                         random_state=21,\n                                         verbose=0,\n                                         warm_start=False,\n                                         class_weight={0 : 1, 1 : 5},\n                                         ccp_alpha=0.,\n                                         max_samples=None\n                                        )\n\n    fit_LR = LR.fit(X_train, y_train)\n    fit_BRF = BRF.fit(X_train, y_train)\n    \n    pred = np.array(0.5 * fit_BRF.predict(X_val) + 0.5 * fit_LR.predict(X_val), dtype = np.uint8)\n    f1 = f1_score(y_val, pred)\n    \n    true = np.array(y_val)\n    pred_LR = np.array(fit_LR.predict_proba(X_val))\n    pred_BRF = np.array(fit_BRF.predict_proba(X_val))\n    pred = 0.5 * pred_BRF + 0.5 * pred_LR\n    val_score = Score.get_score(true, pred)\n    val_scores.append(val_score)\n    \n    print(f'fold: {i}')\n    print(f'f1 val score: {f1}')\n    print(f'w_logloss val score: {val_score}')\n    print()\n    \nprint(f\"mean val score: {np.mean(val_scores)} +- {2 * np.std(val_scores)}\")\n```\n\n# 𝐀𝐧𝐝 𝐠𝐨𝐭 𝐭𝐡𝐞 𝐟𝐨𝐥𝐥𝐨𝐰𝐢𝐧𝐠 𝐫𝐞𝐬𝐮𝐥𝐭𝐬:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11837581%2F4c45ae6d9091e8491ed5ad6553e1a396%2FScreenshot%202023-10-16%20at%2010.44.42.png?generation=1697442310307315&alt=media)\n\n\n\nThis approach showed 0.02 better logloss than the best statistical approach.\nIt seems that the idea can be improved, for example by collecting better data or generating more suitable features.\nThank you for your attention, I look forward to your criticism and suggestions.\nI'm waiting for your questions.\n\n<3\n\n\n"
  }
}