Skip to content

A typo and an attempt at a quality of life change #430

Description

@cotyreh
  1. I found a typo in the getOptionChain.R file:
 # set cleaned up colnames to current output colnames
    cnames <- c(contractsymbol = "ContractID",
                contractsize = "ConractSize",
                currency = "Currency",

On line 22 the renamed contract size column is missing a "t".

  1. I tried and failed to speed up the time it takes to download whole options chains when you use the Exp=NULL argument. My attempt was to add an omit = c() argument to getOptionChain() that would let you specify a list of columns you don't need when you grab the option chain:
getOptionChain.yahoo <- function(Symbols, Exp, ..., session=NULL)
{

  omit <- list(...)$omit   # defined the omit argument here
  
  NewToOld <- function(x, tz = NULL, omit = NULL) {
    if(is.null(x) || length(x) < 1)
      return(NULL)
      
      
# ...

    # convert trade time to exchange timezone
    d$LastTradeTime <- .POSIXct(d$LastTradeTime, tz=tz)
    
    # Omit specified columns if 'omit' is provided
    if(!is.null(omit)) {
      d <- d[ , !(names(d) %in% omit), drop=FALSE]
    }
    
    return(d)
  }
  
  if (is.null(session)) {
    session <- .yahooSession()
  }
  if (!session$can.crumb) {
    stop("Unable to obtain yahoo crumb. If this is being called from a GDPR country, Yahoo requires GDPR consent, which cannot be scripted")
  }
  
# ...

 if(is.null(Exp)) {
      # Return all expiries if Exp = NULL
      out <- lapply(all.expiries, function(e) {
        getOptionChain.yahoo(Symbols, e, .expiry.known=TRUE, session=session, omit=omit) # added omit here
      })
      # Expiry format was "%b %Y", but that's not unique with weeklies. Change
      # format to "%b.%d.%Y" ("%Y-%m-%d wouldn't be good, since names should
      # start with a letter or dot--naming things is hard).
      return(setNames(out, format(all.expiries.posix, "%b.%d.%Y")))
    } else {
      # Ensure data exist for user-provided expiry date(s)
      if(inherits(Exp, "Date"))
        valid.expiries <- as.Date(all.expiries.posix) %in% Exp
      else if(inherits(Exp, "POSIXt"))
        valid.expiries <- all.expiries.posix %in% Exp
      else if(is.character(Exp)) {
        expiry.range <- range(unlist(lapply(Exp, .parseISO8601, tz="UTC")))
        valid.expiries <- all.expiries.posix >= expiry.range[1] &
          all.expiries.posix <= expiry.range[2]
      }
      if(all(!valid.expiries))
        stop("Provided expiry date(s) not found. Available dates are: ",
             paste(as.Date(all.expiries.posix), collapse=", "))
      
      expiry.subset <- all.expiries[valid.expiries]
      if(length(expiry.subset) == 1)
        return(getOptionChain.yahoo(Symbols, expiry.subset, .expiry.known=TRUE, session=session, omit=omit)) # added omit
      else {
        out <- lapply(expiry.subset, function(e) {
          getOptionChain.yahoo(Symbols, e, .expiry.known=TRUE, session=session, omit=omit) # added omit 
        })
        # See comment above regarding the output names
        return(setNames(out, format(all.expiries.posix[valid.expiries], "%b.%d.%Y")))
      }
    }
  }
  

This didn't reduce the time it took to retrieve the data, and I am guessing it's because yahoo sends all the columns and my argument only removes them after the fact. Is there a way to request only a subset of the data from yahoo?

Activity

  1. joshuaulrich commented on Jan 1, 2025

    @joshuaulrich
    Owner

    Thanks for catching the typo!

    Requesting a subset of the data is unlikely to provide a noticeable performance boost. I tested with SPY, which currently has 32 expiry dates. The returned JSON is small (~50Kb per expiry, 2.5Mb total). Requests take an average of 90ms round trip. I'd estimate ~30% of that is network latency (it's 5ms just to my ISP), which is fixed regardless of your internet bandwidth or the amount of data being transferred. Reducing the number of columns requested isn't going to significantly speed up the server's database query because the amount of data is so small. All that said, I don't know how to request a subset of the data anyway. :)

  2. cotyreh commented on Jan 4, 2025

    @cotyreh
    Author

    Interesting. This checks out when I check the timestamps on the data and when I
    clock the run with system.time:

    > system.time(
    +   getOptionChain("^SPX", NULL)
    + )
       user  system elapsed 
      1.627   0.138  13.208 
      
      
    > system.time(
    +   getOptionChain("SPY", NULL)
    + )
       user  system elapsed 
      0.897   0.081   8.051 
    

    I'm not sure where all the extra elapsed time is coming from. It's like I'm getting
    the data right away but then my machine sits around and does nothing for a bit. hmm.

  3. joshuaulrich commented on Jan 4, 2025

    @joshuaulrich
    Owner

    It's the other way around. Your machine is waiting for the response from the Yahoo server. See ?proc.time for a description of the output from system.time().

    For your SPY example, your machine spends 897ms doing actual work, but it takes 8s for the call to finish. That 7s is the time it takes for the request to reach the server, the server to process it, and for the response to come back.

    Here's what it looks like for me:

    R$ system.time(quantmod::getOptionChain("SPY", NULL))
     
       user  system elapsed 
      1.429   0.009   4.145 

    That's 2.716s (4.145 - 1.429) to make 30 requests (90ms per request), which is very reasonable. Yours takes ~240ms per request, which isn't great but isn't unreasonable. My guess is that you're further from the Yahoo servers and/or don't have as good a connection as I do.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions