Our website uses cookies to enhance your experience. By continuing to use our site, or clicking "Continue," you are agreeing to our Cookie Policy | Continue JAMA HomeNew OnlineCurrent IssueFor Authors Podcasts Clinical Reviews Editors' Summary Medical News Author Interviews More Publications JAMA JAMA Network Open JAMA Cardiology JAMA Dermatology JAMA Health Forum JAMA Internal Medicine JAMA Neurology JAMA Oncology JAMA Ophthalmology JAMA Otolaryngology–Head & Neck Surgery JAMA Pediatrics JAMA Psychiatry JAMA Surgery Archives of Neurology & Psychiatry (1919-1959) JN Learning / CMESubscribeJobsInstitutions / LibrariansReprints & Permissions Terms of Use | Privacy Policy | Accessibility Statement 2023 American Medical Association. All Rights Reserved Search All JAMA JAMA Network Open JAMA Cardiology JAMA Dermatology JAMA Forum Archive JAMA Health Forum JAMA Internal Medicine JAMA Neurology JAMA Oncology JAMA Ophthalmology JAMA Otolaryngology–Head & Neck Surgery JAMA Pediatrics JAMA Psychiatry JAMA Surgery Archives of Neurology & Psychiatry Input Search Term Sign In Individual Sign In Sign inCreate an Account Access through your institution Sign In Purchase Options: Buy this article Rent this article Subscribe to the JAMA journal
Ware & Munafo 1 usefully overview the causes and consequences of and possible solutions to the problem of irreproducibility that haunts many currently published results. I am on the same wavelength with almost everything they say. I will discuss here one additional aspect that usually receives less attention in the current discussions about irreproducibility: the need to test with formal experimental studies the multiple solutions that are proposed and to identify measurable outcomes for the benefits but also the potential harms of each proposed intervention. Scientific practices are probably not yet totally broken. Most scientists still mean well, and plain fraud has not overrun science 2. However, scientific practices are fragile and can easily break, as many interested stakeholders pull from all sides. Any effort to improve practices may inadvertently have collateral harms. It behooves all of us who think about how to improve scientific practices to also consider what might go wrong with each intervention. Ideas on how to fix research are just ideas. They may look good, but may not work well in real life. We need to preconceive what might be the potential harms and be prepared to measure them. Moreover, unpredictable harms may also arise and we should be ready to detect them. Like testing new drugs, one wants to be prepared to capture information on known adverse events, but also have sentinel systems in place to capture unanticipated side effects. The more thinking goes into potential benefits and harms of interventions, the better. For example, many suggested improvements are based on reporting checklists (ranging from the methods checklist proposed by Nature 3 to the checklists contemplated by the National Institute of Health (NIH) 4 and the numerous reporting standards checklists for different study designs summarized by EQUATOR (Enhancing the QUAlity and Transparency Of health Research) 5). One may measure improvement by assessing whether adoption of checklists improved the completeness of reporting of requested items. However, this is naive. Improvement in this scale will occur by default. Investigators will comply if funding or acceptance of their paper for publication is dependent upon 27 check-marks. One should also be able to measure potential collateral harms; for example, how many studies report false or inaccurate (and thus misleading) information in their effort to satisfy requirements. Another very good idea is the review of full manuscripts without results before any data collection 6. There are several potential harms here; for example, this may lead to spurious compliance with the pre-specified format in reporting the study, even though the study had to deviate from the original plan for some (good?) reason—many studies, especially with human subjects, have to deviate substantially from the original plan during their conduct. Moreover, there is no guarantee that a pre-approved manuscript can eliminate or even meaningfully reduce data dredging, given that many subtle but influential choices in the analysis plan may not be fully transparent in the pre-approved manuscript 7. Finally, a practice where papers are pre-approved based on having sufficiently high power may force investigators to choose outcomes that have little scientific value, but good power to be assessed, as opposed to outcomes that are scientifically and/or clinically important but only modestly powered 8. A third example is pre-registration. While an excellent idea, much research is simply exploratory and thus cannot be meaningfully pre-registered. Forcing pre-registration for all research will simply force investigators to either abandon exploratory discovery research or to seemingly pre-register research that has already happened, making the literature even more misleading. I list here only three examples of policies which I personally support fervently, and which I have even been among the first to propose. I am still biased to believe that they are worthy of consideration. However, the devil is in the detail. I would loathe seeing a ‘perfect’ scientific literature where everything is pre-registered, all checklist items are checked and papers are written by robotic automata before the research is conducted, but no real progress is made. We need to find ways to improve science without destroying it. We need to find ways to reward excellence: this includes high impact, high quality, reproducibility, sharing culture and eventually the ability to translate information to useful knowledge and applications—and perhaps more 9. We should not compromise for less. None. The Meta-Research Innovation Center at Stanford is supported by a grant from the Laura and John Arnold Foundation.
Despite a plethora of available journals, the most influential papers are extremely concentrated in few journals, especially in fields with high citation density. Existing multidisciplinary journals publish selectively most-cited papers from fields with high citation density.
The effects of small-scale interventions often prove much lower than expected when they are implemented at a large scale. We illustrate the problem and its potential causes using a number of examples from the early childhood intervention literature. We delve deeper by introducing a basic logical framework allowing us to discuss the key factors in assessing whether a program is ready to scale, particularly with regards to uncertainty in the potential outcomes of small-scale interventions. We conclude putting forward a set of concrete recommendations on how to bridge the science of using science and real-life policy.
Selvapatt and colleagues made several points: large numbers of people are infected; birth cohort screening is cost effective; previously treated patients without severe fibrosis are unlikely to progress; sustained virological response improves quality of life; treatment reduces mortality; and newer agents have fewer side effects (last two also alluded to by Matthews and colleagues).1 2 3 The number of patients is not the issue, which is whether …
Abstract OBJECTIVE To provide estimates of the relative risk of COVID-19 death in people <65 years old versus older individuals in the general population, the absolute risk of COVID-19 death at the population level during the first epidemic wave, and the proportion of COVID-19 deaths in non-elderly people without underlying diseases in epicenters of the pandemic. ELIGIBLE DATA Countries and US states with at least 800 COVID-19 deaths as of April 24, 2020 and with information on the number of deaths in people with age <65. Data were available for 11 European countries (Belgium, France, Germany, Ireland, Italy, Netherlands, Portugal, Spain, Sweden, Switzerland, UK), Canada, and 12 US states (California, Connecticut, Florida, Georgia, Illinois, Indiana, Louisiana, Maryland, Massachusetts, Michigan, New Jersey and New York) We also examined available data on COVID-19 deaths in people with age <65 and no underlying diseases. MAIN OUTCOME MEASURES Proportion of COVID-19 deaths in people <65 years old; relative risk of COVID-19 death in people <65 versus ≥65 years old; absolute risk of COVID-19 death in people <65 and in those ≥80 years old in the general population as of May 1, 2020; absolute COVID-19 death risk expressed as equivalent of death risk from driving a motor vehicle. RESULTS Individuals with age <65 account for 4.8-9.3% of all COVID-19 deaths in 10 European countries and Canada, 13.0% in the UK, and 7.8-23.9% in the US locations. People <65 years old had 36- to 84-fold lower risk of COVID-19 death than those ≥65 years old in 10 European countries and Canada and 14- to 56-fold lower risk in UK and US locations. The absolute risk of COVID-19 death as of May 1, 2020 for people <65 years old ranged from 6 (Canada) to 249 per million (New York City). The absolute risk of COVID-19 death for people ≥80 years old ranged from 0.3 (Florida) to 10.6 per thousand (New York). The COVID-19 death risk in people <65 years old during the period of fatalities from the epidemic was equivalent to the death risk from driving between 13 and 101 miles per day for 11 countries and 6 states, and was higher (equivalent to the death risk from driving 143-668 miles per day) for 6 other states and the UK. People <65 years old without underlying predisposing conditions accounted for only 0.7-2.6% of all COVID-19 deaths (data available from France, Italy, Netherlands, Sweden, Georgia, and New York City). CONCLUSIONS People <65 years old have very small risks of COVID-19 death even in pandemic epicenters and deaths for people <65 years without underlying predisposing conditions are remarkably uncommon. Strategies focusing specifically on protecting high-risk elderly individuals should be considered in managing the pandemic.
This paper summarizes the proceedings of an NIAID-sponsored workshop on statistical issues for HIV surrogate endpoints. The workshop brought together statisticians and clinicians in an attempt to shed light on some unresolved issues in the use of HIV laboratory markers (such as HIV RNA and CD4+ cell counts) in the design and analysis of clinical studies and in patient management. Utilizing a debate format, the workshop explored a series of specific questions dealing with the relationship between markers and clinical endpoints, and the choice of endpoints and methods of analysis in clinical studies. This paper provides the position statements from the two debaters on each issue. Consensus conclusions, based on the presentations and discussion, are outlined. While not providing final answers, we hope that these discussions have helped clarify a number of issues, and will stimulate further consideration of some of the highlighted problems. These issues will be critical in the proper assessment and use of future therapies for HIV disease. © 1998 John Wiley & Sons, Ltd.
This paper describes a superscalar processor that combines the best qualities of static and dynamic instruction scheduling to increase the performance of non-numerical applications. The architecture performs all instruction scheduling statically to take advantage of the compiler's ability to efficiently schedule operations across many basic blocks. Since the conditional branches in non-numerical code are highly data dependent, the architecture introduces the concept of boosted instructions, instructions that are committed conditionally upon the result of later branch instructions. Boosting effectively removes the dependencies caused by branches and makes the scheduling of side-effect instructions as simple as those that are side-effect free. For efficiency, boosting is supported in the hardware by shadow structures that temporarily hold the side effects of boosted instructions until the conditional branches that the boosted instructions depend upon are executed. When the branch condition is determined, the buffered side effects are either committed or squashed. The limited static scheduler in our evaluation system shows that a 1.6-times speedup over scalar code is achievable by boosting instructions above only a single conditional branch. This performance is similar to the performance of a pure dynamic scheduler.