{
  "id": 358081,
  "title": "Is Sigmoid Layer necessary in model?",
  "url": "/competitions/rsna-2022-cervical-spine-fracture-detection/discussion/358081",
  "author_name": "Chenjie",
  "post_date": "2022-10-06T14:49:55.657000",
  "votes": 1,
  "comment_count": 7,
  "views": 0,
  "content": "<p>The prediction is between 0 and 1. I add a sigmoid layer after a densenet 3d model, but the model doesn't converge. In <a href=\"https://www.kaggle.com/code/samuelcortinhas/rnsa-3d-model-train-pytorch\" target=\"_blank\">this post</a>, it says sigmoid layer isn't necessary.  But I got negative output if I remove the sigmoid layer. What should I do to build the 3d model?</p>",
  "messages": [
    {
      "id": 1976245,
      "postDate": "2022-10-07T09:08:54.070Z",
      "content": "<p>You should look into what are your labels in the model you are training. Typically you can have fractures in more than one vertebra. In this case your activation function should take this into account, you should avoid any softmax and use sigmoids. Sigmoid not only sets the output between 0 and 1 but also helps to stabilize gradients and avoid divergence. Take into account that after changing activation or loss function learning rate should also be reviewed too many times it is the cause of divergence, as the saying goes: keep calm and lower your learning rate! good luck</p>",
      "rawMarkdown": "You should look into what are your labels in the model you are training. Typically you can have fractures in more than one vertebra. In this case your activation function should take this into account, you should avoid any softmax and use sigmoids. Sigmoid not only sets the output between 0 and 1 but also helps to stabilize gradients and avoid divergence. Take into account that after changing activation or loss function learning rate should also be reviewed too many times it is the cause of divergence, as the saying goes: keep calm and lower your learning rate! good luck",
      "votes": 1,
      "replies": [
        {
          "id": 1977611,
          "postDate": "2022-10-08T06:01:12.773Z",
          "content": "<p>Got it. Thanks for your kind reply. I choose simple sigmoid ,weighted bce loss and 0.0001 learning rate for my baseline model. I got a model with 0.50 loss. It seems everthing works. I am current working on augmentation because my baseline model overfitted.</p>",
          "rawMarkdown": "Got it. Thanks for your kind reply. I choose simple sigmoid ,weighted bce loss and 0.0001 learning rate for my baseline model. I got a model with 0.50 loss. It seems everthing works. I am current working on augmentation because my baseline model overfitted."
        }
      ]
    },
    {
      "id": 1975184,
      "postDate": "2022-10-06T16:22:31.717Z",
      "content": "<p>In Pytorch, it depends on what is your choice of the loss function. If you use <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.BCEWithLogitsLoss.html\" target=\"_blank\">BCEWithLogitsLoss</a>, then no need to use sigmoid because the loss function we run sigmoid inside of it. Otherwise, if you use <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.BCELoss.html\" target=\"_blank\">BCELoss</a>, then you should apply sigmoid to your logits (model's output) before calling the loss.</p>",
      "rawMarkdown": "In Pytorch, it depends on what is your choice of the loss function. If you use [BCEWithLogitsLoss](https://pytorch.org/docs/stable/generated/torch.nn.BCEWithLogitsLoss.html), then no need to use sigmoid because the loss function we run sigmoid inside of it. Otherwise, if you use [BCELoss](https://pytorch.org/docs/stable/generated/torch.nn.BCELoss.html), then you should apply sigmoid to your logits (model's output) before calling the loss.",
      "votes": 1,
      "replies": [
        {
          "id": 1975595,
          "postDate": "2022-10-06T22:07:34.500Z",
          "content": "<p>Thank you for your reply. I really appreciate it。</p>",
          "rawMarkdown": "Thank you for your reply. I really appreciate it。"
        }
      ]
    },
    {
      "id": 1974948,
      "postDate": "2022-10-06T14:49:55.657Z",
      "content": "<p>The prediction is between 0 and 1. I add a sigmoid layer after a densenet 3d model, but the model doesn't converge. In <a href=\"https://www.kaggle.com/code/samuelcortinhas/rnsa-3d-model-train-pytorch\" target=\"_blank\">this post</a>, it says sigmoid layer isn't necessary.  But I got negative output if I remove the sigmoid layer. What should I do to build the 3d model?</p>",
      "rawMarkdown": "The prediction is between 0 and 1. I add a sigmoid layer after a densenet 3d model, but the model doesn't converge. In [this post](https://www.kaggle.com/code/samuelcortinhas/rnsa-3d-model-train-pytorch), it says sigmoid layer isn't necessary.  But I got negative output if I remove the sigmoid layer. What should I do to build the 3d model?",
      "votes": 1
    },
    {
      "id": 1975193,
      "postDate": "2022-10-06T16:24:30.767Z",
      "content": "<p>He is using nn.BCEwithLogitsLoss which apply sigmoid automatically to the logits, so he can directly feed in the logits without sigmoid.., if you are using nn.BCELoss, then you have to do sigmoid before calculating the loss, in both cases, during inference you have to apply sigmoid to make your predictions and make them between 0-1</p>",
      "rawMarkdown": "He is using nn.BCEwithLogitsLoss which apply sigmoid automatically to the logits, so he can directly feed in the logits without sigmoid.., if you are using nn.BCELoss, then you have to do sigmoid before calculating the loss, in both cases, during inference you have to apply sigmoid to make your predictions and make them between 0-1",
      "votes": 2,
      "replies": [
        {
          "id": 1975367,
          "postDate": "2022-10-06T18:23:44.730Z",
          "content": "<p>If you output logits from your model and use BCEwithLogitsLoss, remember to convert logits to probabilities when you make predictions from your inference model.  I would never forget to do that 😨</p>",
          "rawMarkdown": "If you output logits from your model and use BCEwithLogitsLoss, remember to convert logits to probabilities when you make predictions from your inference model.  I would never forget to do that 😨",
          "votes": 1
        },
        {
          "id": 1975591,
          "postDate": "2022-10-06T22:06:37.927Z",
          "content": "<p>Thank you for your reply. It really help me a lot. I would choose nn.BCELoss because I want to do trails on logit methods as in yolov4.</p>",
          "rawMarkdown": "Thank you for your reply. It really help me a lot. I would choose nn.BCELoss because I want to do trails on logit methods as in yolov4."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1976245,
      "author_name": "Cristian Martí",
      "author_url": "",
      "post_date": "2022-10-07T09:08:54.070000",
      "content": "<p>You should look into what are your labels in the model you are training. Typically you can have fractures in more than one vertebra. In this case your activation function should take this into account, you should avoid any softmax and use sigmoids. Sigmoid not only sets the output between 0 and 1 but also helps to stabilize gradients and avoid divergence. Take into account that after changing activation or loss function learning rate should also be reviewed too many times it is the cause of divergence, as the saying goes: keep calm and lower your learning rate! good luck</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1977611,
          "author_name": "Chenjie",
          "author_url": "",
          "post_date": "2022-10-08T06:01:12.773000",
          "content": "<p>Got it. Thanks for your kind reply. I choose simple sigmoid ,weighted bce loss and 0.0001 learning rate for my baseline model. I got a model with 0.50 loss. It seems everthing works. I am current working on augmentation because my baseline model overfitted.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1975184,
      "author_name": "Phat Tran",
      "author_url": "",
      "post_date": "2022-10-06T16:22:31.717000",
      "content": "<p>In Pytorch, it depends on what is your choice of the loss function. If you use <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.BCEWithLogitsLoss.html\" target=\"_blank\">BCEWithLogitsLoss</a>, then no need to use sigmoid because the loss function we run sigmoid inside of it. Otherwise, if you use <a href=\"https://pytorch.org/docs/stable/generated/torch.nn.BCELoss.html\" target=\"_blank\">BCELoss</a>, then you should apply sigmoid to your logits (model's output) before calling the loss.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1975595,
          "author_name": "Chenjie",
          "author_url": "",
          "post_date": "2022-10-06T22:07:34.500000",
          "content": "<p>Thank you for your reply. I really appreciate it。</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1975193,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2022-10-06T16:24:30.767000",
      "content": "<p>He is using nn.BCEwithLogitsLoss which apply sigmoid automatically to the logits, so he can directly feed in the logits without sigmoid.., if you are using nn.BCELoss, then you have to do sigmoid before calculating the loss, in both cases, during inference you have to apply sigmoid to make your predictions and make them between 0-1</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1975367,
          "author_name": "SolverWorld",
          "author_url": "",
          "post_date": "2022-10-06T18:23:44.730000",
          "content": "<p>If you output logits from your model and use BCEwithLogitsLoss, remember to convert logits to probabilities when you make predictions from your inference model.  I would never forget to do that 😨</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1975591,
          "author_name": "Chenjie",
          "author_url": "",
          "post_date": "2022-10-06T22:06:37.927000",
          "content": "<p>Thank you for your reply. It really help me a lot. I would choose nn.BCELoss because I want to do trails on logit methods as in yolov4.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1976245": "You should look into what are your labels in the model you are training. Typically you can have fractures in more than one vertebra. In this case your activation function should take this into account, you should avoid any softmax and use sigmoids. Sigmoid not only sets the output between 0 and 1 but also helps to stabilize gradients and avoid divergence. Take into account that after changing activation or loss function learning rate should also be reviewed too many times it is the cause of divergence, as the saying goes: keep calm and lower your learning rate! good luck",
    "1975184": "In Pytorch, it depends on what is your choice of the loss function. If you use [BCEWithLogitsLoss](https://pytorch.org/docs/stable/generated/torch.nn.BCEWithLogitsLoss.html), then no need to use sigmoid because the loss function we run sigmoid inside of it. Otherwise, if you use [BCELoss](https://pytorch.org/docs/stable/generated/torch.nn.BCELoss.html), then you should apply sigmoid to your logits (model's output) before calling the loss.",
    "1974948": "The prediction is between 0 and 1. I add a sigmoid layer after a densenet 3d model, but the model doesn't converge. In [this post](https://www.kaggle.com/code/samuelcortinhas/rnsa-3d-model-train-pytorch), it says sigmoid layer isn't necessary.  But I got negative output if I remove the sigmoid layer. What should I do to build the 3d model?",
    "1975193": "He is using nn.BCEwithLogitsLoss which apply sigmoid automatically to the logits, so he can directly feed in the logits without sigmoid.., if you are using nn.BCELoss, then you have to do sigmoid before calculating the loss, in both cases, during inference you have to apply sigmoid to make your predictions and make them between 0-1"
  }
}