{
  "id": 450616,
  "title": "submission error",
  "url": "/competitions/UBC-OCEAN/discussion/450616",
  "author_name": "William Green",
  "post_date": "2023-10-25T03:54:49.512000",
  "votes": 3,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I really don't how I am keep getting a submission error.  </p>\n<pre><code> pandas  pd\n collections  Counter\n PIL  Image\n os\n numpy  np\n\n\n ():\n     Image.(tile_path)  img:\n        img_rgb = img.convert()\n         np.array(img_rgb)\n\n\n ():\n    \n    prediction = custom_model.predict(np.expand_dims(tile, axis=))\n     np.argmax(prediction)  \n\n\npreds = []\n\n\ntiles_directory = \n\n\ntile_predictions_dict = {}\n\n\n root, dirs, files  os.walk(tiles_directory):\n     file  files:\n        \n          file.endswith():\n            \n\n        image_id = file.split()[]  \n        tile_path = os.path.join(root, file)\n        tile = load_tile(tile_path)\n        tile_prediction = make_prediction(tile)\n\n        \n         image_id   tile_predictions_dict:\n            tile_predictions_dict[image_id] = []\n        tile_predictions_dict[image_id].append(tile_prediction)\n\n\n image_id, tile_predictions  tile_predictions_dict.items():\n    most_common_prediction = Counter(tile_predictions).most_common()[][]\n    preds.append({: image_id, : most_common_prediction})\n\n\nsample_submission = pd.read_csv()\n\n\npreds_sorted = (preds, key= x: (x[]))\n\n\nlabels_sorted = [item[]  item  preds_sorted]\n\n\nsample_submission[] = labels_sorted\n\n\nlabel_map = {: , : , : , : , :}  \nsample_submission[] = sample_submission[].(label_map)\n\n\nsample_submission.to_csv(, index=)\n\n\n(sample_submission.head())\n</code></pre>",
  "messages": [
    {
      "id": 2497971,
      "postDate": "2023-10-25T03:54:49.513Z",
      "content": "<p>I really don't how I am keep getting a submission error.  </p>\n<pre><code> pandas  pd\n collections  Counter\n PIL  Image\n os\n numpy  np\n\n\n ():\n     Image.(tile_path)  img:\n        img_rgb = img.convert()\n         np.array(img_rgb)\n\n\n ():\n    \n    prediction = custom_model.predict(np.expand_dims(tile, axis=))\n     np.argmax(prediction)  \n\n\npreds = []\n\n\ntiles_directory = \n\n\ntile_predictions_dict = {}\n\n\n root, dirs, files  os.walk(tiles_directory):\n     file  files:\n        \n          file.endswith():\n            \n\n        image_id = file.split()[]  \n        tile_path = os.path.join(root, file)\n        tile = load_tile(tile_path)\n        tile_prediction = make_prediction(tile)\n\n        \n         image_id   tile_predictions_dict:\n            tile_predictions_dict[image_id] = []\n        tile_predictions_dict[image_id].append(tile_prediction)\n\n\n image_id, tile_predictions  tile_predictions_dict.items():\n    most_common_prediction = Counter(tile_predictions).most_common()[][]\n    preds.append({: image_id, : most_common_prediction})\n\n\nsample_submission = pd.read_csv()\n\n\npreds_sorted = (preds, key= x: (x[]))\n\n\nlabels_sorted = [item[]  item  preds_sorted]\n\n\nsample_submission[] = labels_sorted\n\n\nlabel_map = {: , : , : , : , :}  \nsample_submission[] = sample_submission[].(label_map)\n\n\nsample_submission.to_csv(, index=)\n\n\n(sample_submission.head())\n</code></pre>",
      "rawMarkdown": "I really don't how I am keep getting a submission error.  \n\n\n```python\nimport pandas as pd\nfrom collections import Counter\nfrom PIL import Image\nimport os\nimport numpy as np\n\n# Function to load a tile\ndef load_tile(tile_path):\n    with Image.open(tile_path) as img:\n        img_rgb = img.convert('RGB')\n        return np.array(img_rgb)\n\n# Function to make a prediction\ndef make_prediction(tile):\n    # Assume model is your trained model\n    prediction = custom_model.predict(np.expand_dims(tile, axis=0))\n    return np.argmax(prediction)  # Adjust this line as per your model's output\n\n# Placeholder for predictions\npreds = []\n\n# Directory where the tiles are stored\ntiles_directory = '/kaggle/input/test-tiles'\n\n# Dictionary to hold tile predictions for each image_id\ntile_predictions_dict = {}\n\n# Iterate through the directory\nfor root, dirs, files in os.walk(tiles_directory):\n    for file in files:\n        # Skip non-image files\n        if not file.endswith('.png'):\n            continue\n        \n        image_id = file.split('_')[0]  # extract image_id from file name\n        tile_path = os.path.join(root, file)\n        tile = load_tile(tile_path)\n        tile_prediction = make_prediction(tile)\n        \n        # Store tile predictions by image_id\n        if image_id not in tile_predictions_dict:\n            tile_predictions_dict[image_id] = []\n        tile_predictions_dict[image_id].append(tile_prediction)\n\n# Majority vote to get image prediction for each image_id\nfor image_id, tile_predictions in tile_predictions_dict.items():\n    most_common_prediction = Counter(tile_predictions).most_common(1)[0][0]\n    preds.append({'image_id': image_id, 'label': most_common_prediction})\n\n# Read the sample_submission.csv file\nsample_submission = pd.read_csv(\"/kaggle/input/UBC-OCEAN/sample_submission.csv\")\n\n# Ensure the order of image_ids in preds matches the order in sample_submission\npreds_sorted = sorted(preds, key=lambda x: int(x['image_id']))\n\n# Extract only the labels from the sorted preds list\nlabels_sorted = [item['label'] for item in preds_sorted]\n\n# Update the 'label' column of sample_submission\nsample_submission['label'] = labels_sorted\n\n# Map the numerical labels to text labels\nlabel_map = {0: 'CC', 1: 'EC', 2: 'HGSC ', 3: 'LGSC', 4:'MC'}  # etc.\nsample_submission['label'] = sample_submission['label'].map(label_map)\n\n# Save the updated DataFrame to a new CSV file\nsample_submission.to_csv('updated_submission.csv', index=False)\n\n# Display the first few rows of the updated submission DataFrame\nprint(sample_submission.head())\n```\n",
      "votes": 3
    },
    {
      "id": 2502260,
      "postDate": "2023-10-28T04:33:56.530Z",
      "content": "<p>I am also getting this Notebook Out of Memory Error while submitting notebook</p>",
      "rawMarkdown": "I am also getting this Notebook Out of Memory Error while submitting notebook",
      "votes": 1
    },
    {
      "id": 2498068,
      "postDate": "2023-10-25T05:47:05.020Z",
      "content": "<p>The test dataset is too large and an exception occurs when loading the image. i guess load limit is setting. </p>\n<ul>\n<li><p>if you use Image.open()<br>\nPIL.Image.MAX_IMAGE_PIXELS = 933120000</p></li>\n<li><p>or if you use cv2.imread()<br>\nimport os<br>\nos.environ[\"OPENCV_IO_MAX_IMAGE_PIXELS\"] = pow(2,60).<strong>str</strong>()</p></li>\n</ul>\n<p>It would be a good idea to experiment by changing the pixel size. I think the bigger the better.<br>\nand An exception may occur due to another problem, so please keep this in mind.</p>\n<p><strong>Edit</strong> : you need rename sample_submission.to_csv('updated_submission.csv', index=False) <br>\nto sample_submission.to_csv('submission.csv', index=False) </p>",
      "rawMarkdown": "The test dataset is too large and an exception occurs when loading the image. i guess load limit is setting. \n- if you use Image.open()\nPIL.Image.MAX_IMAGE_PIXELS = 933120000\n\n- or if you use cv2.imread()\nimport os\nos.environ[\"OPENCV_IO_MAX_IMAGE_PIXELS\"] = pow(2,60).__str__()\n\nIt would be a good idea to experiment by changing the pixel size. I think the bigger the better.\nand An exception may occur due to another problem, so please keep this in mind.\n\n\n\n**Edit** : you need rename sample_submission.to_csv('updated_submission.csv', index=False) \nto sample_submission.to_csv('submission.csv', index=False) ",
      "votes": 2,
      "replies": [
        {
          "id": 2498834,
          "postDate": "2023-10-25T14:42:11.413Z",
          "content": "<p>The max pixels did not work. Now I get:</p>\n<pre><code>Notebook Threw Exception\nYour notebook hit an unhandled error  rerunning your code. Note that the hidden dataset can be larger/smaller/different than the public dataset See more debugging tips\n</code></pre>\n<p>I added the suggested line:</p>\n<p><code>PIL.Image.MAX_IMAGE_PIXELS = 933120000</code></p>\n<p>This is the frustrating part about kaggle submissions. </p>",
          "rawMarkdown": "The max pixels did not work. Now I get:\n\n```python\nNotebook Threw Exception\nYour notebook hit an unhandled error while rerunning your code. Note that the hidden dataset can be larger/smaller/different than the public dataset See more debugging tips\n```\nI added the suggested line:\n\n`PIL.Image.MAX_IMAGE_PIXELS = 933120000`\n\nThis is the frustrating part about kaggle submissions. ",
          "votes": 1,
          "replies": [
            {
              "id": 2499311,
              "postDate": "2023-10-26T00:59:21.077Z",
              "content": "<p>If you change the image size or use the filtering function, an exception may occur. <br>\nBecause the size of the test set data images is very diverse, the img itself may disappear during processing.</p>",
              "rawMarkdown": "If you change the image size or use the filtering function, an exception may occur. \nBecause the size of the test set data images is very diverse, the img itself may disappear during processing."
            },
            {
              "id": 2499823,
              "postDate": "2023-10-26T09:22:31.627Z",
              "content": "<p>I think we are dealing with a similar error. Did you manage to find a solution?</p>",
              "rawMarkdown": "I think we are dealing with a similar error. Did you manage to find a solution?"
            },
            {
              "id": 2499972,
              "postDate": "2023-10-26T11:48:38.830Z",
              "content": "<p>Not yet, I used up my submissions yesterday trying to figure out. I will try again today. </p>",
              "rawMarkdown": "Not yet, I used up my submissions yesterday trying to figure out. I will try again today. "
            },
            {
              "id": 2500014,
              "postDate": "2023-10-26T12:15:11.257Z",
              "content": "<p>Alright, cool. We managed to get it running with a batch size of 1. It was pretty slow, but the submission went through. Maybe you could try that as well. </p>",
              "rawMarkdown": "Alright, cool. We managed to get it running with a batch size of 1. It was pretty slow, but the submission went through. Maybe you could try that as well. "
            }
          ]
        }
      ]
    },
    {
      "id": 2503451,
      "postDate": "2023-10-29T05:40:13.623Z",
      "content": "<p>The errors I experienced were 1. when resizing the image after reading it, and 2. when filtering (filtering according to the ratio of black and white parts) was applied after reading the image.</p>\n<p>In case 1, resize was applied only when the image size exceeded a certain standard, and in case 2, no error occurred when the filter was applied when the black and white ratio was less than 0.8 (an error occurred when applied at 0.7).</p>\n<p>If you have applied similar processing to mine, it would be a good idea to modify that part.<br>\nThis is because the test set has an image set of too small a size, so there may not be images that meet the criteria when resized or filtered.</p>",
      "rawMarkdown": "The errors I experienced were 1. when resizing the image after reading it, and 2. when filtering (filtering according to the ratio of black and white parts) was applied after reading the image.\n\nIn case 1, resize was applied only when the image size exceeded a certain standard, and in case 2, no error occurred when the filter was applied when the black and white ratio was less than 0.8 (an error occurred when applied at 0.7).\n\nIf you have applied similar processing to mine, it would be a good idea to modify that part.\nThis is because the test set has an image set of too small a size, so there may not be images that meet the criteria when resized or filtered."
    }
  ],
  "comments": [
    {
      "id": 2502260,
      "author_name": "sunil thite",
      "author_url": "",
      "post_date": "2023-10-28T04:33:56.530000",
      "content": "<p>I am also getting this Notebook Out of Memory Error while submitting notebook</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2498068,
      "author_name": "Trappist",
      "author_url": "",
      "post_date": "2023-10-25T05:47:05.020000",
      "content": "<p>The test dataset is too large and an exception occurs when loading the image. i guess load limit is setting. </p>\n<ul>\n<li><p>if you use Image.open()<br>\nPIL.Image.MAX_IMAGE_PIXELS = 933120000</p></li>\n<li><p>or if you use cv2.imread()<br>\nimport os<br>\nos.environ[\"OPENCV_IO_MAX_IMAGE_PIXELS\"] = pow(2,60).<strong>str</strong>()</p></li>\n</ul>\n<p>It would be a good idea to experiment by changing the pixel size. I think the bigger the better.<br>\nand An exception may occur due to another problem, so please keep this in mind.</p>\n<p><strong>Edit</strong> : you need rename sample_submission.to_csv('updated_submission.csv', index=False) <br>\nto sample_submission.to_csv('submission.csv', index=False) </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2498834,
          "author_name": "William Green",
          "author_url": "",
          "post_date": "2023-10-25T14:42:11.413000",
          "content": "<p>The max pixels did not work. Now I get:</p>\n<pre><code>Notebook Threw Exception\nYour notebook hit an unhandled error  rerunning your code. Note that the hidden dataset can be larger/smaller/different than the public dataset See more debugging tips\n</code></pre>\n<p>I added the suggested line:</p>\n<p><code>PIL.Image.MAX_IMAGE_PIXELS = 933120000</code></p>\n<p>This is the frustrating part about kaggle submissions. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2499311,
              "author_name": "Trappist",
              "author_url": "",
              "post_date": "2023-10-26T00:59:21.077000",
              "content": "<p>If you change the image size or use the filtering function, an exception may occur. <br>\nBecause the size of the test set data images is very diverse, the img itself may disappear during processing.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2499823,
              "author_name": "Elias Leijonmarck",
              "author_url": "",
              "post_date": "2023-10-26T09:22:31.627000",
              "content": "<p>I think we are dealing with a similar error. Did you manage to find a solution?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2499972,
              "author_name": "William Green",
              "author_url": "",
              "post_date": "2023-10-26T11:48:38.830000",
              "content": "<p>Not yet, I used up my submissions yesterday trying to figure out. I will try again today. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2500014,
              "author_name": "Elias Leijonmarck",
              "author_url": "",
              "post_date": "2023-10-26T12:15:11.257000",
              "content": "<p>Alright, cool. We managed to get it running with a batch size of 1. It was pretty slow, but the submission went through. Maybe you could try that as well. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2503451,
      "author_name": "Trappist",
      "author_url": "",
      "post_date": "2023-10-29T05:40:13.623000",
      "content": "<p>The errors I experienced were 1. when resizing the image after reading it, and 2. when filtering (filtering according to the ratio of black and white parts) was applied after reading the image.</p>\n<p>In case 1, resize was applied only when the image size exceeded a certain standard, and in case 2, no error occurred when the filter was applied when the black and white ratio was less than 0.8 (an error occurred when applied at 0.7).</p>\n<p>If you have applied similar processing to mine, it would be a good idea to modify that part.<br>\nThis is because the test set has an image set of too small a size, so there may not be images that meet the criteria when resized or filtered.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2497971": "I really don't how I am keep getting a submission error.  \n\n\n```python\nimport pandas as pd\nfrom collections import Counter\nfrom PIL import Image\nimport os\nimport numpy as np\n\n# Function to load a tile\ndef load_tile(tile_path):\n    with Image.open(tile_path) as img:\n        img_rgb = img.convert('RGB')\n        return np.array(img_rgb)\n\n# Function to make a prediction\ndef make_prediction(tile):\n    # Assume model is your trained model\n    prediction = custom_model.predict(np.expand_dims(tile, axis=0))\n    return np.argmax(prediction)  # Adjust this line as per your model's output\n\n# Placeholder for predictions\npreds = []\n\n# Directory where the tiles are stored\ntiles_directory = '/kaggle/input/test-tiles'\n\n# Dictionary to hold tile predictions for each image_id\ntile_predictions_dict = {}\n\n# Iterate through the directory\nfor root, dirs, files in os.walk(tiles_directory):\n    for file in files:\n        # Skip non-image files\n        if not file.endswith('.png'):\n            continue\n        \n        image_id = file.split('_')[0]  # extract image_id from file name\n        tile_path = os.path.join(root, file)\n        tile = load_tile(tile_path)\n        tile_prediction = make_prediction(tile)\n        \n        # Store tile predictions by image_id\n        if image_id not in tile_predictions_dict:\n            tile_predictions_dict[image_id] = []\n        tile_predictions_dict[image_id].append(tile_prediction)\n\n# Majority vote to get image prediction for each image_id\nfor image_id, tile_predictions in tile_predictions_dict.items():\n    most_common_prediction = Counter(tile_predictions).most_common(1)[0][0]\n    preds.append({'image_id': image_id, 'label': most_common_prediction})\n\n# Read the sample_submission.csv file\nsample_submission = pd.read_csv(\"/kaggle/input/UBC-OCEAN/sample_submission.csv\")\n\n# Ensure the order of image_ids in preds matches the order in sample_submission\npreds_sorted = sorted(preds, key=lambda x: int(x['image_id']))\n\n# Extract only the labels from the sorted preds list\nlabels_sorted = [item['label'] for item in preds_sorted]\n\n# Update the 'label' column of sample_submission\nsample_submission['label'] = labels_sorted\n\n# Map the numerical labels to text labels\nlabel_map = {0: 'CC', 1: 'EC', 2: 'HGSC ', 3: 'LGSC', 4:'MC'}  # etc.\nsample_submission['label'] = sample_submission['label'].map(label_map)\n\n# Save the updated DataFrame to a new CSV file\nsample_submission.to_csv('updated_submission.csv', index=False)\n\n# Display the first few rows of the updated submission DataFrame\nprint(sample_submission.head())\n```\n",
    "2502260": "I am also getting this Notebook Out of Memory Error while submitting notebook",
    "2498068": "The test dataset is too large and an exception occurs when loading the image. i guess load limit is setting. \n- if you use Image.open()\nPIL.Image.MAX_IMAGE_PIXELS = 933120000\n\n- or if you use cv2.imread()\nimport os\nos.environ[\"OPENCV_IO_MAX_IMAGE_PIXELS\"] = pow(2,60).__str__()\n\nIt would be a good idea to experiment by changing the pixel size. I think the bigger the better.\nand An exception may occur due to another problem, so please keep this in mind.\n\n\n\n**Edit** : you need rename sample_submission.to_csv('updated_submission.csv', index=False) \nto sample_submission.to_csv('submission.csv', index=False) ",
    "2503451": "The errors I experienced were 1. when resizing the image after reading it, and 2. when filtering (filtering according to the ratio of black and white parts) was applied after reading the image.\n\nIn case 1, resize was applied only when the image size exceeded a certain standard, and in case 2, no error occurred when the filter was applied when the black and white ratio was less than 0.8 (an error occurred when applied at 0.7).\n\nIf you have applied similar processing to mine, it would be a good idea to modify that part.\nThis is because the test set has an image set of too small a size, so there may not be images that meet the criteria when resized or filtered."
  }
}