{
  "id": 426666,
  "title": "Importance of Pretraining? ",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/426666",
  "author_name": "Blaine Heffron",
  "post_date": "2023-07-24T14:29:19.440000",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi there,</p>\n<p>I'm curious about the importance of using a pretrained model for this problem. The paper talks about it and the UNet implementation someone posted that scores ~.62 also uses pretrained weights. </p>\n<p>I'm fairly new to image segmentation and I wanted to learn a fast framework for TPU training so I implemented DeepLabV3+ in jax. However, my model caps out around .25 for the dice coefficient. My suspicion is that I have an error in the implementation but I'm wondering if maybe I just need to pretrain the model on imagenet or some such dataset? Could pretraining really be that important? My intuition would be that pretraining should only help you get a model trained up faster, it shouldnt actually help with the final accuracy of the model. But again I'm fairly new to this so maybe my intuition is off. </p>",
  "messages": [
    {
      "id": 2358062,
      "postDate": "2023-07-25T09:22:16.353Z",
      "content": "<p>Not 100% related, but <a href=\"https://magazine.sebastianraschka.com/p/accelerating-pytorch-model-training\" target=\"_blank\">here</a> you'll find a discussion on the impact of using pre-trained models to train ViTs. It does indeed decrease the training time but it also boost considerably the model performance.</p>",
      "rawMarkdown": "Not 100% related, but [here](https://magazine.sebastianraschka.com/p/accelerating-pytorch-model-training) you'll find a discussion on the impact of using pre-trained models to train ViTs. It does indeed decrease the training time but it also boost considerably the model performance.",
      "votes": 1,
      "replies": [
        {
          "id": 2365898,
          "postDate": "2023-07-30T14:41:35.553Z",
          "content": "<p>Thanks Fabien! This is very helpful. </p>",
          "rawMarkdown": "Thanks Fabien! This is very helpful. "
        }
      ]
    },
    {
      "id": 2357079,
      "postDate": "2023-07-24T14:29:19.440Z",
      "content": "<p>Hi there,</p>\n<p>I'm curious about the importance of using a pretrained model for this problem. The paper talks about it and the UNet implementation someone posted that scores ~.62 also uses pretrained weights. </p>\n<p>I'm fairly new to image segmentation and I wanted to learn a fast framework for TPU training so I implemented DeepLabV3+ in jax. However, my model caps out around .25 for the dice coefficient. My suspicion is that I have an error in the implementation but I'm wondering if maybe I just need to pretrain the model on imagenet or some such dataset? Could pretraining really be that important? My intuition would be that pretraining should only help you get a model trained up faster, it shouldnt actually help with the final accuracy of the model. But again I'm fairly new to this so maybe my intuition is off. </p>",
      "rawMarkdown": "Hi there,\n\nI'm curious about the importance of using a pretrained model for this problem. The paper talks about it and the UNet implementation someone posted that scores ~.62 also uses pretrained weights. \n\nI'm fairly new to image segmentation and I wanted to learn a fast framework for TPU training so I implemented DeepLabV3+ in jax. However, my model caps out around .25 for the dice coefficient. My suspicion is that I have an error in the implementation but I'm wondering if maybe I just need to pretrain the model on imagenet or some such dataset? Could pretraining really be that important? My intuition would be that pretraining should only help you get a model trained up faster, it shouldnt actually help with the final accuracy of the model. But again I'm fairly new to this so maybe my intuition is off. \n\n",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2358062,
      "author_name": "FabienDaniel",
      "author_url": "",
      "post_date": "2023-07-25T09:22:16.353000",
      "content": "<p>Not 100% related, but <a href=\"https://magazine.sebastianraschka.com/p/accelerating-pytorch-model-training\" target=\"_blank\">here</a> you'll find a discussion on the impact of using pre-trained models to train ViTs. It does indeed decrease the training time but it also boost considerably the model performance.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2365898,
          "author_name": "Blaine Heffron",
          "author_url": "",
          "post_date": "2023-07-30T14:41:35.553000",
          "content": "<p>Thanks Fabien! This is very helpful. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2358062": "Not 100% related, but [here](https://magazine.sebastianraschka.com/p/accelerating-pytorch-model-training) you'll find a discussion on the impact of using pre-trained models to train ViTs. It does indeed decrease the training time but it also boost considerably the model performance.",
    "2357079": "Hi there,\n\nI'm curious about the importance of using a pretrained model for this problem. The paper talks about it and the UNet implementation someone posted that scores ~.62 also uses pretrained weights. \n\nI'm fairly new to image segmentation and I wanted to learn a fast framework for TPU training so I implemented DeepLabV3+ in jax. However, my model caps out around .25 for the dice coefficient. My suspicion is that I have an error in the implementation but I'm wondering if maybe I just need to pretrain the model on imagenet or some such dataset? Could pretraining really be that important? My intuition would be that pretraining should only help you get a model trained up faster, it shouldnt actually help with the final accuracy of the model. But again I'm fairly new to this so maybe my intuition is off. \n\n"
  }
}